DocumentCode
323570
Title
A discriminant measure for model complexity adaptation
Author
Bahl, L.R. ; Padmanabhan, M.
Author_Institution
IBM Thomas J. Watson Res. Center, Yorktown Heights, NY, USA
Volume
1
fYear
1998
fDate
12-15 May 1998
Firstpage
453
Abstract
We present a discriminant measure that can be used to determine the model complexity in a speech recognition system. In the speech recognition process, given a test feature vector the conditional probability of the feature vector has to be obtained for several allophone (sub-phonetic units) classes using a Gaussian-mixture density model for each class. The Gaussian-mixture models are constructed from the training data belonging to the allophone classes, and the number of mixture components that are required to adequately model the PDF of each class is determined by using some simple rule of thumb-for instance the number of components has to be sufficient to model the data reasonably well but not so many as to overmodel the data. A typical example of the choice of the number is to make it proportional to the number of data samples. However, such methods may result in models that are sub-optimal as far as classification accuracy is concerned. We present a new discriminant measure that can be used to determine in an objective fashion, the number of Gaussians required to best model the PDF of an allophone class. We also present the results of experiments showing the improvement in recognition performance when the number of mixture components is chosen based on the discriminant measure as opposed to the rule of thumb. These results are presented both for the speaker-independent and speaker-adapted case
Keywords
Gaussian processes; computational complexity; feature extraction; probability; speech recognition; Gaussian-mixture density model; PDF; allophone classes; classification accuracy; conditional probability; data samples; discriminant measure; experiments; mixture components; model complexity adaptation; recognition performance; rule of thumb; speaker-adapted recognition; speaker-independent recognition; speech recognition system; sub-phonetic units; test feature vector; training data; Adaptation model; Data mining; Gaussian processes; Parametric statistics; Probability; Speech processing; Speech recognition; Testing; Thumb; Training data;
fLanguage
English
Publisher
ieee
Conference_Titel
Acoustics, Speech and Signal Processing, 1998. Proceedings of the 1998 IEEE International Conference on
Conference_Location
Seattle, WA
ISSN
1520-6149
Print_ISBN
0-7803-4428-6
Type
conf
DOI
10.1109/ICASSP.1998.674465
Filename
674465
Link To Document