Personalized Spectral and Prosody Conversion Using Frame-Based Codeword Distribution and Adaptive CRF

Author

Huang, Yi-Chin ; Wu, Chung-Hsien ; Chao, Yu-Ting

Author_Institution

Dept. of Comput. Sci. & Inf. Eng., Nat. Cheng Kung Univ., Tainan, Taiwan

Volume

21

Issue

1

fYear

2013

fDate

Jan. 2013

Firstpage

51

Lastpage

62

Abstract

This study proposes a voice conversion-based approach to personalized text-to-speech (TTS) synthesis. The conversion functions, trained using a small parallel corpus with source and target speech data, can impose the voice characteristics of a target speaker on an existing synthesizer. Frame alignment between a pair of sentences in the parallel corpus is generally used for training voice conversion functions. However, with incorrect alignment, the resultant conversion functions may generate unacceptable conversion results. Traditional frame alignment using minimal spectral distance between the frame-based feature vectors of the source and the target phone sequences can be imprecise because the voice properties of the source and target phones inherently differ. In the proposed method, feature vectors of the parallel corpus are transformed into codewords in an eigenspace. A more precise frame alignment can be obtained by integrating the codeword occurrence distributions into distance estimation. In addition to the spectral property, a prosodic word/phrase boundary prediction model was constructed using an adaptive conditional random field (CRF) to generate personalized prosodic information. Objective and subjective tests were conducted to evaluate the performance of the proposed approach. The experimental results showed that the proposed voice conversion method, based on distribution-based alignment and prosodic word boundary detection, can improve the speech quality and speaker similarity of the converted speech. Compared to other methods, the evaluation results verified the improved performance of the proposed method.

Keywords

speech coding; speech synthesis; TTS synthesis; adaptive CRF; codeword occurrence distributions; conditional random field; distance estimation; frame-based codeword distribution; personalized spectral; personalized text-to-speech synthesis; prosodic word boundary detection; prosody conversion; speech synthesis; synthesizer; target phone sequences; target speech data; voice conversion-based approach; Adaptation models; Hidden Markov models; Predictive models; Speech; Speech synthesis; Synthesizers; Training; Conditional random field; frame alignment; principal component analysis; prosodic boundary; voice conversion;

fLanguage

English

Journal_Title

Audio, Speech, and Language Processing, IEEE Transactions on

Publisher

ieee

ISSN

1558-7916

Type

jour

DOI

10.1109/TASL.2012.2213247

Filename

6269060