DocumentCode
2311431
Title
Audio-visual speaker identification using coupled hidden Markov models
Author
Fu, Tieyan ; Liu, Xiao Xing ; Liang, Lu Hong ; Pi, Xiaobo ; Nefian, Ara V.
Author_Institution
Dept. of Comput. Sci. & Technol., Tsinghua Univ., Beijing, China
Volume
3
fYear
2003
fDate
14-17 Sept. 2003
Abstract
In this paper, we investigate the use of the coupled hidden Markov models (CHMM) for the task of audio-visual text dependent speaker identification. Our system determines the identity of the user from a temporal sequence of audio and visual observations obtained from the acoustic speech and the shape of the mouth, respectively. The multi modal observation sequences are then modeled using a set of CHMMs, one for each phoneme-viseme pair and for each person in the database. The use of CHMMs in our system is justified by the capacity of this model to describe the natural audio and visual state asynchrony as well as their conditional dependency over time. To train a CHMM we first train a speaker independent model using expectation-maximization (EM), and then we build a speaker dependent model using maximum a posteriori (MAP) training. Experimental results on XM2VTS database show that our system improves the accuracy of audio-only or video-only speaker identification at all levels of acoustic signal-to-noise ratio (SNR) from 0 to 30 dB.
Keywords
audio-visual systems; hidden Markov models; maximum likelihood estimation; optimisation; speaker recognition; 0 to 30 dB; MAP; XM2VTS database; audio temporal sequence; audio-only speaker identification; audio-visual speaker identification; coupled HMM; coupled hidden Markov models; expectation-maximization; maximum a posteriori; multimodal observation sequences; phoneme-viseme pair; signal-to-noise ratio; speaker dependent model; speaker independent model; video-only speaker identification; visual observation; Computer science; Hidden Markov models; Loudspeakers; Mouth; Noise robustness; Shape; Signal to noise ratio; Speech; Spine; Visual databases;
fLanguage
English
Publisher
ieee
Conference_Titel
Image Processing, 2003. ICIP 2003. Proceedings. 2003 International Conference on
ISSN
1522-4880
Print_ISBN
0-7803-7750-8
Type
conf
DOI
10.1109/ICIP.2003.1247173
Filename
1247173
Link To Document