Title :
A unified framework for clustering heterogeneous Web objects
Author :
Zeng, Hua-Jun ; Chen, Zheng ; Ma, Wei-Ying
Author_Institution :
Microsoft Res. Asia, Beijing, China
Abstract :
We introduce a novel framework for clustering Web data which is often heterogeneous in nature. As most existing methods often integrate heterogeneous data into a unified feature space, their flexibilities to explore and adjust contributing effects from different heterogeneous information are compromised. In contrast, our framework enables separate clustering of homogeneous data in the entire process based on their respective features, and a layered structure with link information is used to iteratively project and propagate the clustered results between layers until it converges. Our experimental results show that such a scheme not only effectively overcomes the problem of data sparseness caused by the high dimensional link space but also improves the clustering accuracy significantly. We achieve 19% and 41% performance increases when clustering Web-pages and users based on a semi-synthetic Web log. Finally, we show a real clustering result based on UC Berkeley´s Web log.
Keywords :
information resources; information retrieval; Web data; Web log; Web pages; clustering accuracy; data sparseness; heterogeneous Web object clustering; high dimensional link space; layered structure; link information; Asia; Books; Collaboration; Electrical capacitance tomography; Information filtering; Information filters; Privacy; Sea measurements; Tellurium; Web sites;
Conference_Titel :
Web Information Systems Engineering, 2002. WISE 2002. Proceedings of the Third International Conference on
Print_ISBN :
0-7695-1766-8
DOI :
10.1109/WISE.2002.1181653