DocumentCode
2792053
Title
Unit selection speech synthesis using multiple speech units at non-adjacent segments for prosody and waveform generation
Author
Tamura, Masatsune ; Braunschweiler, Norbert ; Kagoshima, Takehiko ; Akamine, Masami
Author_Institution
Corp. R&D Center, Toshiba Corp., Kawasaki, Japan
fYear
2010
fDate
14-19 March 2010
Firstpage
4802
Lastpage
4805
Abstract
In this paper, we propose a speech synthesis method that combines a natural waveform concatenation based speech synthesis method and our baseline plural unit selection and fusion method. Two main features of the proposed method are (i) prosody regeneration from selected speech units and (ii) using multiple speech units at non-adjacent segments. The nonadjacent segments is the segment that the previous or following speech units in the optimum speech unit sequence are not adjacent in the database. By using the prosody of selected speech units, the original prosodic expressions and sounds of recorded speech are retained, while discontinuities are reduced by using multiple speech units at non-adjacent segments. MOS evaluations showed that the proposed method provides a clear improvement against the conventional unit selection method and our baseline method.
Keywords
linguistics; speech synthesis; waveform analysis; MOS evaluation; baseline plural unit selection; fusion method; natural waveform concatenation; non adjacent segments; prosody regeneration; speech synthesis method; unit selection; Cost function; Degradation; Europe; Frequency; Fusion power generation; Laboratories; Natural languages; Research and development; Spatial databases; Speech synthesis; concatenative speech synthesis; prosody generation; unit fusion; unit selection;
fLanguage
English
Publisher
ieee
Conference_Titel
Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on
Conference_Location
Dallas, TX
ISSN
1520-6149
Print_ISBN
978-1-4244-4295-9
Electronic_ISBN
1520-6149
Type
conf
DOI
10.1109/ICASSP.2010.5495151
Filename
5495151
Link To Document