DocumentCode :
2311431
Title :
Audio-visual speaker identification using coupled hidden Markov models
Author :
Fu, Tieyan ; Liu, Xiao Xing ; Liang, Lu Hong ; Pi, Xiaobo ; Nefian, Ara V.
Author_Institution :
Dept. of Comput. Sci. & Technol., Tsinghua Univ., Beijing, China
Volume :
3
fYear :
2003
fDate :
14-17 Sept. 2003
Abstract :
In this paper, we investigate the use of the coupled hidden Markov models (CHMM) for the task of audio-visual text dependent speaker identification. Our system determines the identity of the user from a temporal sequence of audio and visual observations obtained from the acoustic speech and the shape of the mouth, respectively. The multi modal observation sequences are then modeled using a set of CHMMs, one for each phoneme-viseme pair and for each person in the database. The use of CHMMs in our system is justified by the capacity of this model to describe the natural audio and visual state asynchrony as well as their conditional dependency over time. To train a CHMM we first train a speaker independent model using expectation-maximization (EM), and then we build a speaker dependent model using maximum a posteriori (MAP) training. Experimental results on XM2VTS database show that our system improves the accuracy of audio-only or video-only speaker identification at all levels of acoustic signal-to-noise ratio (SNR) from 0 to 30 dB.
Keywords :
audio-visual systems; hidden Markov models; maximum likelihood estimation; optimisation; speaker recognition; 0 to 30 dB; MAP; XM2VTS database; audio temporal sequence; audio-only speaker identification; audio-visual speaker identification; coupled HMM; coupled hidden Markov models; expectation-maximization; maximum a posteriori; multimodal observation sequences; phoneme-viseme pair; signal-to-noise ratio; speaker dependent model; speaker independent model; video-only speaker identification; visual observation; Computer science; Hidden Markov models; Loudspeakers; Mouth; Noise robustness; Shape; Signal to noise ratio; Speech; Spine; Visual databases;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Image Processing, 2003. ICIP 2003. Proceedings. 2003 International Conference on
ISSN :
1522-4880
Print_ISBN :
0-7803-7750-8
Type :
conf
DOI :
10.1109/ICIP.2003.1247173
Filename :
1247173
Link To Document :
بازگشت