• DocumentCode
    323570
  • Title

    A discriminant measure for model complexity adaptation

  • Author

    Bahl, L.R. ; Padmanabhan, M.

  • Author_Institution
    IBM Thomas J. Watson Res. Center, Yorktown Heights, NY, USA
  • Volume
    1
  • fYear
    1998
  • fDate
    12-15 May 1998
  • Firstpage
    453
  • Abstract
    We present a discriminant measure that can be used to determine the model complexity in a speech recognition system. In the speech recognition process, given a test feature vector the conditional probability of the feature vector has to be obtained for several allophone (sub-phonetic units) classes using a Gaussian-mixture density model for each class. The Gaussian-mixture models are constructed from the training data belonging to the allophone classes, and the number of mixture components that are required to adequately model the PDF of each class is determined by using some simple rule of thumb-for instance the number of components has to be sufficient to model the data reasonably well but not so many as to overmodel the data. A typical example of the choice of the number is to make it proportional to the number of data samples. However, such methods may result in models that are sub-optimal as far as classification accuracy is concerned. We present a new discriminant measure that can be used to determine in an objective fashion, the number of Gaussians required to best model the PDF of an allophone class. We also present the results of experiments showing the improvement in recognition performance when the number of mixture components is chosen based on the discriminant measure as opposed to the rule of thumb. These results are presented both for the speaker-independent and speaker-adapted case
  • Keywords
    Gaussian processes; computational complexity; feature extraction; probability; speech recognition; Gaussian-mixture density model; PDF; allophone classes; classification accuracy; conditional probability; data samples; discriminant measure; experiments; mixture components; model complexity adaptation; recognition performance; rule of thumb; speaker-adapted recognition; speaker-independent recognition; speech recognition system; sub-phonetic units; test feature vector; training data; Adaptation model; Data mining; Gaussian processes; Parametric statistics; Probability; Speech processing; Speech recognition; Testing; Thumb; Training data;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Acoustics, Speech and Signal Processing, 1998. Proceedings of the 1998 IEEE International Conference on
  • Conference_Location
    Seattle, WA
  • ISSN
    1520-6149
  • Print_ISBN
    0-7803-4428-6
  • Type

    conf

  • DOI
    10.1109/ICASSP.1998.674465
  • Filename
    674465