DocumentCode :
2228203
Title :
Clustering for Web information hierarchy mining
Author :
Kao, Hung-Yu ; Chen, Ming-Syan ; Ho, Jan-Ming
Author_Institution :
Dept. of Electr. Eng., Nat. Taiwan Univ., Taipei, Taiwan
fYear :
2003
fDate :
13-17 Oct. 2003
Firstpage :
698
Lastpage :
701
Abstract :
Benefiting from the growth of techniques of dynamic page generation, the amount and the complexity of Web pages increase explosively. The structures of Web pages which are dynamically generated by the same templates are thus similar to one another and are usually assembled by a set of fundamental information clusters These neighboring information clusters usually represent the similar semantics and form a larger cluster with the more generalized information. The hierarchical structure generated by information clusters in a bottom-up manner is called the information hierarchy of a page. We study the problem of mining the information hierarchies of pages in Web sites to recognize the information distribution of pages within the multilevel, multigranularity configurations. Explicitly, we propose an information clustering system that applies a top-down information centroid searching algorithm and a multigranularity centroid converging process on the document object model (DOM) trees of pages to build the information hierarchies of pages. Experiments on several real news Web sites show the high precision and recall rates of the proposed method on determining information clusters of pages and also validate its practical applicability to real Web sites.
Keywords :
Internet; Web sites; data mining; document handling; information retrieval; statistical analysis; DOM tree; Web information hierarchy mining; Web site; document object model; dynamic Web page generation; information clustering system; information page distribution; multigranularity centroid converging process; top-down information centroid searching algorithm; Assembly; Clustering algorithms; Data mining; Electronic mail; HTML; Information science; Joining processes; Merging; Scalability; Web pages;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Web Intelligence, 2003. WI 2003. Proceedings. IEEE/WIC International Conference on
Print_ISBN :
0-7695-1932-6
Type :
conf
DOI :
10.1109/WI.2003.1241299
Filename :
1241299
Link To Document :
بازگشت