DocumentCode
2404051
Title
A Topic Model of Observing Chinese Characters
Author
Zhang, Yunkai ; Qin, Zengchang
Author_Institution
Coll. of Software, Beihang Univ., Beijing, China
Volume
2
fYear
2010
fDate
26-28 Aug. 2010
Firstpage
7
Lastpage
10
Abstract
The Topic Models are a class of hierarchical statistical models for analyzing document collections and it has become one of the most used techniques in Natural Language Processing in the recent years. It assumes that each document could be expressed as a mixture of topics and each topic could be characterized by a distribution over words. In previous research, like in English language, Topic Models for Chinese Language use the words as observing data. In this research, we demonstrated the effectiveness of using Chinese characters as the basic units of observing data. The comparisons with those models based on Chinese words and English words are presented.
Keywords
document handling; natural language processing; statistical analysis; Chinese characters; Chinese words; English words; document collections; hierarchical statistical models; natural language processing; topic models; Accuracy; Analytical models; Biological system modeling; Computational modeling; Probabilistic logic; Semantics; Training;
fLanguage
English
Publisher
ieee
Conference_Titel
Intelligent Human-Machine Systems and Cybernetics (IHMSC), 2010 2nd International Conference on
Conference_Location
Nanjing, Jiangsu
Print_ISBN
978-1-4244-7869-9
Type
conf
DOI
10.1109/IHMSC.2010.99
Filename
5591057
Link To Document