Title :
A Co-occurrence Based Hierarchical Method for Clustering Web Search Results
Author :
Zhang, Yun ; Feng, BoQin
Author_Institution :
Sch. of Electron. & Inf. Eng., Xian Jiaotong Univ., Xian
Abstract :
This study proposes a novel method to group and organize search results. We apply statistical techniques to term co-occurrence information in a corpus to retrieve bi-grams firstly, and then combine bi-grams into n-grams. After eliminating redundant n-grams, the remaining ones are ranked and selected as cluster labels. Base clusters are constructed according to these cluster labels and then agglomerated into higher-level clusters. We refer to the proposed algorithm as CoHC (co-occurrence based hierarchical clustering). we compare CoHC with three other search results clustering (SRC) algorithms: suffix tree clustering (STC), Lingo, and Vivisimo. We also analyze the properties of cluster labels produced by different SRC algorithms. The experimental results show that our method outperforms the other three SRC algorithms, and is helpful to the user for browsing and locating the results of interest.
Keywords :
Internet; information retrieval; pattern clustering; statistical analysis; Lingo; Vivisimo; Web search result clustering; cooccurrence based hierarchical clustering; statistical technique; suffix tree clustering; Algorithm design and analysis; Clustering algorithms; Data mining; Information retrieval; Intelligent agent; Internet; Scattering; Search engines; Singular value decomposition; Web search; CoHC (Co-occurrence based hierarchical clustering); clustering evaluation; labeling-then-clustering approach; search results clustering;
Conference_Titel :
Web Intelligence and Intelligent Agent Technology, 2008. WI-IAT '08. IEEE/WIC/ACM International Conference on
Conference_Location :
Sydney, NSW
Print_ISBN :
978-0-7695-3496-1
DOI :
10.1109/WIIAT.2008.35