DocumentCode :
2430171
Title :
Similarity model and term association for document categorization
Author :
Kou, Huaizhong ; Gardarin, Georges
Author_Institution :
PRISM Lab., Univ. of Versailles, France
fYear :
2002
fDate :
2-6 Sept. 2002
Firstpage :
256
Lastpage :
260
Abstract :
Both Euclidean distance- and cosine-based similarity models are widely used for measures of document similarity in information retrieval and document categorization. These two similarity models are based on the assumption that term vectors are orthogonal. But this assumption is not true. Term associations are ignored in such similarity models. In the document categorization context, we analyze the properties of term-document space, term-category space and category-document space. Then, without the assumption of term independence, we propose a new mathematical model to estimate the association between terms and define an ε-similarity model of documents. Here we make best use of existing category membership represented by the corpus as much as possible, and the objective is to improve categorization performance. Experiments have been done with a k-NN classifier over the Reuters-5178 corpus. The empirical results show that utilization of term association can improve the effectiveness of the categorization system and the ε-similarity model outperforms those without term association.
Keywords :
category theory; content-based retrieval; learning (artificial intelligence); pattern classification; text analysis; ϵ-similarity model; Euclidean distance-based similarity model; Reuters-5178 corpus; categorization performance; category membership; category-document space; cosine-based similarity model; document categorization; document similarity; information retrieval; k nearest neighbors algorithm; k-NN classifier; term association; term-category space; term-document space; Conferences; Databases; Euclidean distance; Expert systems; Humans; Information retrieval; Mathematical model; Nearest neighbor searches; Support vector machines; Text categorization;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Database and Expert Systems Applications, 2002. Proceedings. 13th International Workshop on
ISSN :
1529-4188
Print_ISBN :
0-7695-1668-8
Type :
conf
DOI :
10.1109/DEXA.2002.1045908
Filename :
1045908
Link To Document :
بازگشت