• DocumentCode
    3527677
  • Title

    Robust discriminative keyword spotting for emotionally colored spontaneous speech using bidirectional LSTM networks

  • Author

    Wöllmer, Martin ; Eyben, Florian ; Keshet, Joseph ; Graves, Alex ; Schuller, Björn ; Rigoll, Gerhard

  • Author_Institution
    Inst. for Human-Machine Commun., Tech. Univ. Munchen, Munich
  • fYear
    2009
  • fDate
    19-24 April 2009
  • Firstpage
    3949
  • Lastpage
    3952
  • Abstract
    In this paper we propose a new technique for robust keyword spotting that uses bidirectional long short-term memory (BLSTM) recurrent neural nets to incorporate contextual information in speech decoding. Our approach overcomes the drawbacks of generative HMM modeling by applying a discriminative learning procedure that non-linearly maps speech features into an abstract vector space. By incorporating the outputs of a BLSTM network into the speech features, it is able to make use of past and future context for phoneme predictions. The robustness of the approach is evaluated on a keyword spotting task using the HUMAINE sensitive artificial listener (SAL) database, which contains accented, spontaneous, and emotionally colored speech. The test is particularly stringent because the system is not trained on the SAL database, but only on the TIMIT corpus of read speech. We show that our method prevails over a discriminative keyword spotter without BLSTM-enhanced feature functions, which in turn has been proven to outperform HMM-based techniques.
  • Keywords
    decoding; hidden Markov models; speech coding; HUMAINE sensitive artificial listener database; abstract vector space; bidirectional long short-term memory recurrent neural nets; hidden Markov model; robust discriminative keyword spotting; speech decoding; Computer science; Context; Hidden Markov models; Man machine systems; Neural networks; Recurrent neural networks; Robustness; Spatial databases; Speech enhancement; Speech recognition; Recurrent neural networks; Robustness; Speech recognition;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Acoustics, Speech and Signal Processing, 2009. ICASSP 2009. IEEE International Conference on
  • Conference_Location
    Taipei
  • ISSN
    1520-6149
  • Print_ISBN
    978-1-4244-2353-8
  • Electronic_ISBN
    1520-6149
  • Type

    conf

  • DOI
    10.1109/ICASSP.2009.4960492
  • Filename
    4960492