DocumentCode
3184325
Title
Single Stream DBN Model Based Triphone for Continuous Speech Recognition
Author
Lv, Guoyun ; Jiang, Dongmei ; Zhao, Rongchun
fYear
2007
fDate
10-12 Dec. 2007
Firstpage
240
Lastpage
245
Abstract
In this paper, based on an single stream word- phone Dynamic Bayesian Network (WP-DBN) model and an single stream word-phone-state DBN (WPS- DBN) model proposed by Guoyun et al [8], to more accurately capture the variations in real continuous speech spectra, context-dependent triphone models are considered, two single stream DBN models, word- triphone DBN (WT-DBN) model and word-triphone- state DBN (WTS-DBN) model, are proposed for continuous speech recognition. Simultaneously, decision tree-based state tying clustering method is used to maintain the balance between model complexity and their corresponding available training data. Essentially, WTS-DBN model is a triphone model whose recognition modeling units are triphones, and simulates a conventional triphone Hidden Markov Model (HMM). Recognition experiments are done on continuous speech database, and results show that WTS-DBN model has the best performance in speech recognition rate. In clean speech environment, comparing with triphone HMM, WPS-DBN model and WT-DBN model, the improvements of 20.53%, 7.52% and 40.77% are obtained for WTS-DBN model respectively in speech recognition rate.
Keywords
Automatic speech recognition; Bayesian methods; Context modeling; Databases; Decision trees; Hidden Markov models; Speech analysis; Speech recognition; Streaming media; Vocabulary;
fLanguage
English
Publisher
ieee
Conference_Titel
Multimedia Workshops, 2007. ISMW '07. Ninth IEEE International Symposium on
Conference_Location
Taichung, Taiwan
Print_ISBN
9780-7695-3084-0
Type
conf
DOI
10.1109/ISM.Workshops.2007.48
Filename
4475977
Link To Document