• DocumentCode
    179892
  • Title

    Speaker Adaptive Training using Deep Neural Networks

  • Author

    Ochiai, Toshihiko ; Matsuda, Shodai ; Xugang Lu ; Hori, Chiori ; Katagiri, Souichi

  • Author_Institution
    Spoken Language Commun. Lab., Nat. Inst. of Inf. & Commun. TechnolSpoken Language Communication Laboratory, Kyoto, Japan
  • fYear
    2014
  • fDate
    4-9 May 2014
  • Firstpage
    6349
  • Lastpage
    6353
  • Abstract
    Among many speaker adaptation embodiments, Speaker Adaptive Training (SAT) has been successfully applied to a standard Hidden-Markov-Model (HMM) speech recognizer, whose state is associated with Gaussian Mixture Models (GMMs). On the other hand, recent studies on Speaker-Independent (SI) recognizer development have reported that a new type of HMM speech recognizer, which replaces GMMs with Deep Neural Networks (DNNs), outperforms GMM-HMM recognizers. Along these two lines, it is natural to conceive of further improvement to a preset DNN-HMM recognizer by employing SAT. In this paper, we propose a novel training scheme that applies SAT to a SI DNN-HMM recognizer. We then implement the SAT scheme by allocating a Speaker-Dependent (SD) module to one of the intermediate layers of a seven-layer DNN, and elaborate its utility over TED Talks corpus data. Experiment results show that our speaker-adapted SAT-based DNN-HMM recognizer reduces the word error rate by 8.4% more than that of a baseline SI DNN-HMM recognizer, and (regardless of the SD module allocation) outperforms the conventional speaker adaptation scheme. The results also show that the inner layers of DNN are more suitable for the SD module than the outer layers.
  • Keywords
    Gaussian processes; hidden Markov models; learning (artificial intelligence); neural nets; speaker recognition; Gaussian mixture models; HMM speech recognizer; SAT; SD module; baseline SI DNN-HMM recognizer; deep neural networks; hidden-Markov-model; seven-layer DNN; speaker adaptation scheme; speaker adaptive training; speaker-dependent module; speaker-independent recognizer development; training scheme; Acoustics; Hidden Markov models; Silicon; Speech; Speech recognition; Training; Vectors; Deep Neural Network; Speaker Adaptative Training;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on
  • Conference_Location
    Florence
  • Type

    conf

  • DOI
    10.1109/ICASSP.2014.6854826
  • Filename
    6854826