مرکز منطقه ای اطلاع رساني علوم و فناوري - Creating synthetic voices for children by adapting adult average voice using stacked transformations and VTLN

DocumentCode :

3164001

Title :

Creating synthetic voices for children by adapting adult average voice using stacked transformations and VTLN

Author :

Karhila, Reima ; Sanand, D.R. ; Kurimo, Mikko ; Smit, Peter

Author_Institution :

Adaptive Inf. Res. Center, Aalto Univ., Aalto, Finland

fYear :

2012

fDate :

25-30 March 2012

Firstpage :

4501

Lastpage :

4504

Abstract :

This paper describes experiments in creating personalised children´s voices for HMM-based synthesis by adapting either an adult or child average voice. The adult average voice is trained from a large adult speech database, whereas the child average voice is trained using a small database of children´s speech. Here we present the idea to use stacked transformations for creating synthetic child voices, where the child average voice is first created from the adult average voice through speaker adaptation using all the pooled speech data from multiple children and then adding child specific speaker adaptation on top of it. VTLN is applied to speech synthesis to see whether it helps the speaker adaptation when only a small amount of adaptation data is available. The listening test results show that the stacked transformations significantly improve speaker adaptation for small amounts of data, but the additional benefit provided by VTLN is not yet clear.

Keywords :

hidden Markov models; speech recognition; speech synthesis; VTLN; adaptation data; adult average voice; adult speech database; child average voice; child specific speaker adaptation; hidden Markov models; speech synthesis; stacked transformations; synthetic child voices; vocal track length normalization; Adaptation models; Data models; Hidden Markov models; Speech; Speech synthesis; Training data; Transforms; Adaptation; Child Speech; Speech synthesis; Stacked Transformations; VTLN;

fLanguage :

English

Publisher :

ieee

Conference_Titel :

Acoustics, Speech and Signal Processing (ICASSP), 2012 IEEE International Conference on

Conference_Location :

Kyoto

ISSN :

1520-6149

Print_ISBN :

978-1-4673-0045-2

Electronic_ISBN :

1520-6149

Type :

conf

DOI :

10.1109/ICASSP.2012.6288918

Filename :

6288918

Link To Document :

https://search.ricest.ac.ir/dl/search/defaultta.aspx?DTC=49&DC=3164001