• DocumentCode
    3211555
  • Title

    A unified framework for clustering heterogeneous Web objects

  • Author

    Zeng, Hua-Jun ; Chen, Zheng ; Ma, Wei-Ying

  • Author_Institution
    Microsoft Res. Asia, Beijing, China
  • fYear
    2002
  • fDate
    12-14 Dec. 2002
  • Firstpage
    161
  • Lastpage
    170
  • Abstract
    We introduce a novel framework for clustering Web data which is often heterogeneous in nature. As most existing methods often integrate heterogeneous data into a unified feature space, their flexibilities to explore and adjust contributing effects from different heterogeneous information are compromised. In contrast, our framework enables separate clustering of homogeneous data in the entire process based on their respective features, and a layered structure with link information is used to iteratively project and propagate the clustered results between layers until it converges. Our experimental results show that such a scheme not only effectively overcomes the problem of data sparseness caused by the high dimensional link space but also improves the clustering accuracy significantly. We achieve 19% and 41% performance increases when clustering Web-pages and users based on a semi-synthetic Web log. Finally, we show a real clustering result based on UC Berkeley´s Web log.
  • Keywords
    information resources; information retrieval; Web data; Web log; Web pages; clustering accuracy; data sparseness; heterogeneous Web object clustering; high dimensional link space; layered structure; link information; Asia; Books; Collaboration; Electrical capacitance tomography; Information filtering; Information filters; Privacy; Sea measurements; Tellurium; Web sites;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Web Information Systems Engineering, 2002. WISE 2002. Proceedings of the Third International Conference on
  • Print_ISBN
    0-7695-1766-8
  • Type

    conf

  • DOI
    10.1109/WISE.2002.1181653
  • Filename
    1181653