Title :
A Unified Trajectory Tiling Approach to High Quality Speech Rendering
Author :
Yao Qian ; Soong, Frank K. ; Zhi-Jie Yan
Author_Institution :
Microsoft Res. Asia, Beijing, China
Abstract :
It is technically challenging to make a machine talk as naturally as a human so as to facilitate “frictionless” interactions between machine and human. We propose a trajectory tiling-based approach to high-quality speech rendering, where speech parameter trajectories, extracted from natural, processed, or synthesized speech, are used to guide the search for the best sequence of waveform “tiles” stored in a pre-recorded speech database. We test the proposed unified algorithm in both Text-To-Speech (TTS) synthesis and cross-lingual voice transformation applications. Experimental results show that the proposed trajectory tiling approach can render speech which is both natural and highly intelligible. The perceived high quality of rendered speech is also confirmed in both objective and subjective evaluations.
Keywords :
speech processing; speech synthesis; cross-lingual voice transformation; frictionless interactions; high quality speech rendering; machine talk; natural speech; pre-recorded speech database; processed speech; speech parameter trajectory; synthesized speech; text-to-speech synthesis; unified trajectory tiling; waveform tiles; Hidden Markov models; Rendering (computer graphics); Speech; Speech processing; Training data; Trajectory; Cross-lingual; speech synthesis; trajectory tiling; voice transformation;
Journal_Title :
Audio, Speech, and Language Processing, IEEE Transactions on
DOI :
10.1109/TASL.2012.2221460