• DocumentCode
    2024075
  • Title

    Classifying Words: A Syllables-Based Model

  • Author

    Warintarawej, P. ; Laurent, A. ; Pompidor, P. ; Cassanas, A. ; Laurent, B.

  • Author_Institution
    LIRMM, Univ. Montpellier 2, Montpellier, France
  • fYear
    2011
  • fDate
    Aug. 29 2011-Sept. 2 2011
  • Firstpage
    208
  • Lastpage
    212
  • Abstract
    Text classification has been extensively studied by linguists and computer scientists. However, there are very few works on classification of words into classes or concepts (e.g. thesaurus). In this paper, we consider this topic, especially in the context of the classification of names like brand names or neologisms. The challenge is thus to provide automated tools to analyze new names by classifying them into concepts. Then, for example, a naming company customer can be informed about which concept a new name is closest to. As we argue that a word can belong to several concepts, we propose to consider the top-k classification approach. Moreover, we rely on syllables to build the classification model. The word corpus is collected from French thesaurus. All labeled-words are separated into syllables. Feature selection techniques are used to select discriminative syllables. We use a syllables frequency (SF) and mutual information (MI) performing with Naive Bayes classifier and K-nearest neighbor (KNN). Instead of selecting only one class, the model select top-k classes ranking them by a classifier score. The result shows the top-k classification model helps to analyze a new word by showing that it can be related to more than one concept. Moreover, the set of discriminative syllables can be used to explain the classification results which makes the results more meaningful.
  • Keywords
    Bayes methods; pattern classification; text analysis; word processing; French thesaurus; K-nearest neighbor; automated tools; classifier score; computer scientists; discriminative syllables; feature selection; labeled words; linguists; naive Bayes classifier; naming company customer; neologism; syllables frequency; syllables-based model; text classification; top-k classes; top-k classification model; word corpus; words classification; Accuracy; Machine learning; Mutual information; Pain; Testing; Thesauri; Training; Discriminative Features; Feature Selection; Syllables; Text Classification; Words Classification;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Database and Expert Systems Applications (DEXA), 2011 22nd International Workshop on
  • Conference_Location
    Toulouse
  • ISSN
    1529-4188
  • Print_ISBN
    978-1-4577-0982-1
  • Type

    conf

  • DOI
    10.1109/DEXA.2011.21
  • Filename
    6059819