مرکز منطقه ای اطلاع رساني علوم و فناوري - Turkish Document Classification Using Shorter Roots

DocumentCode :

3356463

Title :

Turkish Document Classification Using Shorter Roots

Author :

Çataltepe, Zehra ; Turan, Yakup ; Kesgin, Fatih

Author_Institution :

Istanbul Teknik Univ., Istanbul, Turkey

fYear :

2007

fDate :

11-13 June 2007

Firstpage :

Lastpage :

Abstract :

Stemming is one of commonly used pre-processing steps in document categorization. Especially when fast and accurate classification of a lot of documents is needed, it is important to have as small number of and as small length roots as possible. This would not only reduce the time it takes to train and test classifiers but also would reduce the storage requirements for each document. In this study, we analyze the performance of classifiers when the longest or shortest roots found by a stemmer are used. We also analyze the effect of using only the consonants in the roots. We use two document data sets, obtained from Milliyet newspaper and Wikipedia to analyze classification accuracy of classifiers when roots obtained under these four conditions are used. We also analyze the classification accuracy when only the first 4, 3 or 2 letters or consonants are used from the roots. Using smaller roots results in smaller number of TF-IDF vectors. Especially for small sized TF-IDF vectors, using only consonants in the roots gives better performance than using all letters in the roots.

Keywords :

classification; document handling; natural language processing; pattern classification; text analysis; vectors; TF-IDF vectors; Turkish document classification; Wikipedia; classification accuracy; document categorization; document data sets; stemming; storage requirements; Frequency; Performance analysis; Testing; Wikipedia;

fLanguage :

English

Publisher :

ieee

Conference_Titel :

Signal Processing and Communications Applications, 2007. SIU 2007. IEEE 15th

Conference_Location :

Eskisehir

Print_ISBN :

1-4244-0719-2

Electronic_ISBN :

1-4244-0720-6

Type :

conf

DOI :

10.1109/SIU.2007.4298786

Filename :

4298786

Link To Document :

https://search.ricest.ac.ir/dl/search/defaultta.aspx?DTC=49&DC=3356463