• DocumentCode
    3421494
  • Title

    Cross-lingual speaker adaptation for HMM-based speech synthesis considering differences between language-dependent average voices

  • Author

    Peng, Xianglin ; Oura, Keiichiro ; Nankaku, Yoshihiko ; Tokuda, Keiichi

  • Author_Institution
    Dept. of Comput. Sci., Nagoya Inst. of Technol., Nagoya, Japan
  • fYear
    2010
  • fDate
    24-28 Oct. 2010
  • Firstpage
    605
  • Lastpage
    608
  • Abstract
    This paper proposes an improved cross-lingual speaker adaptation technique with considering the differences between language-dependent average voices in a Speech-to-Speech Translation system. A state mapping based method had been introduced for cross-lingual speaker adaptation in HMM-based speech synthesis. In this method, the transforms estimated from the input language are applied to average voice models of the output language according to the state mapping information. However, the differences between average voices in the input and output language may degrade the adaptation performance. To reduce the differences, we apply a global linear transform to output average voice models, which minimizes the symmetric Kullback-Leibler divergence between two average voice models. From the experimental results, our approach could not obtain a better result than the original state mapping based method. This is because the global transform affects not only speaker characteristics but also language identity in acoustic features, and this degrades the synthetic speech quality. Therefore, it becomes clear that a technique which separate speaker and language identities is required.
  • Keywords
    hidden Markov models; natural language processing; speech synthesis; HMM-based speech synthesis; cross-lingual speaker adaptation technique; global linear transform; language-dependent average voice model; speaker characteristics; speech-to-speech translation system; state mapping based method; symmetric Kullback-Leibler divergence; synthetic speech quality; Adaptation model; Data models; Databases; Hidden Markov models; Speech; Speech synthesis; Transforms; HMM; average voice; cross-lingual speaker adaptation; speech synthesis;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Signal Processing (ICSP), 2010 IEEE 10th International Conference on
  • Conference_Location
    Beijing
  • Print_ISBN
    978-1-4244-5897-4
  • Type

    conf

  • DOI
    10.1109/ICOSP.2010.5656849
  • Filename
    5656849