Title :
A new method of finding provisional boundaries of “bunsetsu” using 2nd-order Markov model
Author :
Araki, Tetsuo ; IKEHARA, Satoru ; Tuchihase, J.
Author_Institution :
Fac. of Eng., Fukui Univ., Japan
Abstract :
As Japanese sentences are usually written using thousand kinds of characters especially “kanji” characters, it is not easy to input them into computer files. There has been much research on the method which translates the non-segmented “kana” sentences into the “kanji-kana” sentences. However, the amount of computer memory required for the translating processing explodes in many times, because the number of the combinations of candidates for “kanji-kana” words grows rapidly in proportion to the increasing of the length of the sentence. The memory explosion can be prevented if a sentence is separated into “bunsetsu” This paper proposes a new method of finding provisional boundaries of “bunsetsu” of non-segmented “kana” sentences using 2nd-order Markov chain probabilities. “Relevance factor” P and “Recall factor” R for provisional boundaries of “bunsetsu” determined by this method, were evaluated by experiment using the statistical data for 70 issues of a daily Japanese newspaper
Keywords :
Markov processes; probability; speech recognition; 2nd-order Markov chain probabilities; 2nd-order Markov model; Japanese sentences; Recall factor; Relevance factor; bunsetsu; daily Japanese newspaper; kanji; kanji-kana sentences; nonsegmented kana sentences; provisional boundaries; statistical data; Communication networks; Computer networks; Explosions; Humans; Information analysis; Information systems; Intelligent networks; Intelligent robots; Intelligent structures; Laboratories;
Conference_Titel :
Robot and Human Communication, 1993. Proceedings., 2nd IEEE International Workshop on
Conference_Location :
Tokyo
Print_ISBN :
0-7803-1407-7
DOI :
10.1109/ROMAN.1993.367738