Title of article :
A comparative study of TF*IDF, LSI and multi-words for text classification
Author/Authors :
Zhang، نويسنده , , Ya-Wen and Yoshida، نويسنده , , Taketoshi and Tang، نويسنده , , Xijin، نويسنده ,
Issue Information :
روزنامه با شماره پیاپی سال 2011
Pages :
8
From page :
2758
To page :
2765
Abstract :
One of the main themes in text mining is text representation, which is fundamental and indispensable for text-based intellegent information processing. Generally, text representation inludes two tasks: indexing and weighting. This paper has comparatively studied TF*IDF, LSI and multi-word for text representation. We used a Chinese and an English document collection to respectively evaluate the three methods in information retreival and text categorization. Experimental results have demonstrated that in text categorization, LSI has better performance than other methods in both document collections. Also, LSI has produced the best performance in retrieving English documents. This outcome has shown that LSI has both favorable semantic and statistical quality and is different with the claim that LSI can not produce discriminative power for indexing.
Keywords :
Text classification , Text Categorization , Text representation , information retrieval , lSI , TF*IDF , Multi-word
Journal title :
Expert Systems with Applications
Serial Year :
2011
Journal title :
Expert Systems with Applications
Record number :
2348921
Link To Document :
بازگشت