DocumentCode :
2429766
Title :
A clustering approach for XML linked documents
Author :
Catania, Barbara ; Maddalena, Anna
Author_Institution :
Univ. of Genova, Italy
fYear :
2002
fDate :
2-6 Sept. 2002
Firstpage :
121
Lastpage :
125
Abstract :
Clustering algorithms for hypertext documents consider not only the document content but also the links existing between them. All the similarity functions proposed in the literature assume that just one type of link exists between documents, with a unique semantic meaning. With the rapid diffusion of XML documents, a specific language, called XLink, has been proposed to specify inside XML documents different types of links. Each type of link forces a different degree of similarity between the documents on which it is defined, thus we claim it must influence in a different way the computation of distance values. In this paper, after presenting a graph-based formalization of the hypertexts we consider, we introduce a distance function, based on both the number and the type of the links connecting documents. Some preliminary experimental results on clustering algorithms based on the proposed function conclude the paper.
Keywords :
graph theory; hypermedia markup languages; information resources; pattern clustering; WWW; XLink; XML linked documents; clustering approach; distance function; distance values; document similarity; graph-based formalization; hypertext documents; hypertexts; Clustering algorithms; Conferences; Databases; Expert systems; Joining processes; Merging; Scattering; Unsupervised learning; Visualization; XML;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Database and Expert Systems Applications, 2002. Proceedings. 13th International Workshop on
ISSN :
1529-4188
Print_ISBN :
0-7695-1668-8
Type :
conf
DOI :
10.1109/DEXA.2002.1045887
Filename :
1045887
Link To Document :
بازگشت