Title :
A clustering approach for XML linked documents
Author :
Catania, Barbara ; Maddalena, Anna
Author_Institution :
Univ. of Genova, Italy
Abstract :
Clustering algorithms for hypertext documents consider not only the document content but also the links existing between them. All the similarity functions proposed in the literature assume that just one type of link exists between documents, with a unique semantic meaning. With the rapid diffusion of XML documents, a specific language, called XLink, has been proposed to specify inside XML documents different types of links. Each type of link forces a different degree of similarity between the documents on which it is defined, thus we claim it must influence in a different way the computation of distance values. In this paper, after presenting a graph-based formalization of the hypertexts we consider, we introduce a distance function, based on both the number and the type of the links connecting documents. Some preliminary experimental results on clustering algorithms based on the proposed function conclude the paper.
Keywords :
graph theory; hypermedia markup languages; information resources; pattern clustering; WWW; XLink; XML linked documents; clustering approach; distance function; distance values; document similarity; graph-based formalization; hypertext documents; hypertexts; Clustering algorithms; Conferences; Databases; Expert systems; Joining processes; Merging; Scattering; Unsupervised learning; Visualization; XML;
Conference_Titel :
Database and Expert Systems Applications, 2002. Proceedings. 13th International Workshop on
Print_ISBN :
0-7695-1668-8
DOI :
10.1109/DEXA.2002.1045887