DocumentCode :
2276582
Title :
A Co-occurrence Based Hierarchical Method for Clustering Web Search Results
Author :
Zhang, Yun ; Feng, BoQin
Author_Institution :
Sch. of Electron. & Inf. Eng., Xian Jiaotong Univ., Xian
Volume :
1
fYear :
2008
fDate :
9-12 Dec. 2008
Firstpage :
407
Lastpage :
410
Abstract :
This study proposes a novel method to group and organize search results. We apply statistical techniques to term co-occurrence information in a corpus to retrieve bi-grams firstly, and then combine bi-grams into n-grams. After eliminating redundant n-grams, the remaining ones are ranked and selected as cluster labels. Base clusters are constructed according to these cluster labels and then agglomerated into higher-level clusters. We refer to the proposed algorithm as CoHC (co-occurrence based hierarchical clustering). we compare CoHC with three other search results clustering (SRC) algorithms: suffix tree clustering (STC), Lingo, and Vivisimo. We also analyze the properties of cluster labels produced by different SRC algorithms. The experimental results show that our method outperforms the other three SRC algorithms, and is helpful to the user for browsing and locating the results of interest.
Keywords :
Internet; information retrieval; pattern clustering; statistical analysis; Lingo; Vivisimo; Web search result clustering; cooccurrence based hierarchical clustering; statistical technique; suffix tree clustering; Algorithm design and analysis; Clustering algorithms; Data mining; Information retrieval; Intelligent agent; Internet; Scattering; Search engines; Singular value decomposition; Web search; CoHC (Co-occurrence based hierarchical clustering); clustering evaluation; labeling-then-clustering approach; search results clustering;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Web Intelligence and Intelligent Agent Technology, 2008. WI-IAT '08. IEEE/WIC/ACM International Conference on
Conference_Location :
Sydney, NSW
Print_ISBN :
978-0-7695-3496-1
Type :
conf
DOI :
10.1109/WIIAT.2008.35
Filename :
4740483
Link To Document :
بازگشت