• DocumentCode
    1281526
  • Title

    Structured Discriminative Models For Speech Recognition: An Overview

  • Author

    Gales, Mark ; Watanabe, Shinji ; Fosler-Lussier, Eric

  • Author_Institution
    Eng. Dept., Cambridge Univ., Cambridge, UK
  • Volume
    29
  • Issue
    6
  • fYear
    2012
  • Firstpage
    70
  • Lastpage
    81
  • Abstract
    Automatic speech recognition (ASR) systems classify structured sequence data, where the label sequences (sentences) must be inferred from the observation sequences (the acoustic waveform). The sequential nature of the task is one of the reasons why generative classifiers, based on combining hidden Markov model (HMM) acoustic models and N-gram language models using Bayes rule, have become the dominant technology used in ASR. Conversely, machine learning and natural language processing (NLP) research areas are increasingly dominated by discriminative approaches, where the class posteriors are directly modeled. This article describes recent work in the area of structured discriminative models for ASR. To handle continuous, variable length observation sequences, the approaches applied to NLP tasks must be modified. This article discusses a variety of approaches for applying structured discriminative models to ASR, both from the current literature and possible future approaches. We concentrate on structured models themselves, the descriptive features of observations commonly used within the models, and various options for optimizing the parameters of the model.
  • Keywords
    hidden Markov models; learning (artificial intelligence); natural language processing; speech recognition; ASR systems; Bayes´ rule; HMM acoustic models; N-gram language models; NLP research; NLP tasks; acoustic waveform; automatic speech recognition systems; hidden Markov model acoustic models; machine learning; natural language processing; structured discriminative models; structured sequence data; Acoustics; Adaptation models; Automatic speech recognition; Hidden Markov models; Modeling; Speech recognition; Training;
  • fLanguage
    English
  • Journal_Title
    Signal Processing Magazine, IEEE
  • Publisher
    ieee
  • ISSN
    1053-5888
  • Type

    jour

  • DOI
    10.1109/MSP.2012.2207140
  • Filename
    6296527