• DocumentCode
    2773282
  • Title

    Stemming Versus Light Stemming as Feature Selection Techniques for Arabic Text Categorization

  • Author

    Duwairi, Rehab ; Al-Refai, Mohammad ; Khasawneh, Natheer

  • Author_Institution
    Qatar Univ., Doha
  • fYear
    2007
  • fDate
    18-20 Nov. 2007
  • Firstpage
    446
  • Lastpage
    450
  • Abstract
    This paper compares and contrasts two feature selection techniques when applied to Arabic corpus; in particular; stemming, and light stemming were employed. With stemming, words are reduced to their stems. With light stemming, words are reduced to their light stems. Stemming is aggressive in the sense that it reduces words to their 3-letters roots. This affects the semantics as several words with different meanings might have the same root. Light stemming, by comparison, removes frequently used prefixes and suffixes in Arabic words. Light stemming doesn´t produce the root and therefore doesn´t affect the semantics of words; it maps several words, which have the same meaning to a common syntactical form. The effectiveness of above two feature selection techniques was assessed in a text categorization exercise for Arabic corpus. This corpus consists of 15000 documents that fall into three categories. The K-nearest neighbors (KNN) classifier was used in this work. Several experiments were carried out using two different representations of the same corpus; the first version uses stem- vectors; and the second uses light stem-vectors as representatives of documents. These two representations were assessed in terms of size, time and accuracy. The light stem representation was superior in terms of classifier accuracy when compared with stemming.
  • Keywords
    text analysis; Arabic corpus; Arabic text categorization; K-nearest neighbors classifier; feature selection; light stemming; Classification tree analysis; Computer science; Information filtering; Information filters; Information management; Information resources; Internet; Resource management; Sorting; Text categorization; Arabic language; K-nearest neighbors classifier; feature selection; light-stemming; stemming; text categorization;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Innovations in Information Technology, 2007. IIT '07. 4th International Conference on
  • Conference_Location
    Dubai
  • Print_ISBN
    978-1-4244-1840-4
  • Electronic_ISBN
    978-1-4244-1841-1
  • Type

    conf

  • DOI
    10.1109/IIT.2007.4430403
  • Filename
    4430403