• DocumentCode
    2404051
  • Title

    A Topic Model of Observing Chinese Characters

  • Author

    Zhang, Yunkai ; Qin, Zengchang

  • Author_Institution
    Coll. of Software, Beihang Univ., Beijing, China
  • Volume
    2
  • fYear
    2010
  • fDate
    26-28 Aug. 2010
  • Firstpage
    7
  • Lastpage
    10
  • Abstract
    The Topic Models are a class of hierarchical statistical models for analyzing document collections and it has become one of the most used techniques in Natural Language Processing in the recent years. It assumes that each document could be expressed as a mixture of topics and each topic could be characterized by a distribution over words. In previous research, like in English language, Topic Models for Chinese Language use the words as observing data. In this research, we demonstrated the effectiveness of using Chinese characters as the basic units of observing data. The comparisons with those models based on Chinese words and English words are presented.
  • Keywords
    document handling; natural language processing; statistical analysis; Chinese characters; Chinese words; English words; document collections; hierarchical statistical models; natural language processing; topic models; Accuracy; Analytical models; Biological system modeling; Computational modeling; Probabilistic logic; Semantics; Training;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Intelligent Human-Machine Systems and Cybernetics (IHMSC), 2010 2nd International Conference on
  • Conference_Location
    Nanjing, Jiangsu
  • Print_ISBN
    978-1-4244-7869-9
  • Type

    conf

  • DOI
    10.1109/IHMSC.2010.99
  • Filename
    5591057