DocumentCode
2768553
Title
Monolingual and crosslingual comparison of tandem features derived from articulatory and phone MLPS
Author
Çetin, Özgür ; Magimai-Doss, Mathew ; Livescu, Karen ; Kantor, Arthur ; King, Simon ; Bartels, Chris ; Frankel, Joe
Author_Institution
Yahoo! Inc., Santa Clara
fYear
2007
fDate
9-13 Dec. 2007
Firstpage
36
Lastpage
41
Abstract
The features derived from posteriors of a multilayer perceptron (MLP), known as tandem features, have proven to be very effective for automatic speech recognition. Most tandem features to date have relied on MLPs trained for phone classification. We recently showed on a relatively small data set that MLPs trained for articulatory feature classification can be equally effective. In this paper, we provide a similar comparison using MLPs trained on a much larger data set -2000 hours of English conversational telephone speech. We also explore how portable phone-and articulatory feature-based tandem features are in an entirely different language - Mandarin - without any retraining. We find that while the phone-based features perform slightly better than AF-based features in the matched-language condition, they perform significantly better in the cross-language condition. However, in the cross-language condition, neither approach is as effective as the tandem features extracted from an MLP trained on a relatively small amount of in-domain data. Beyond feature concatenation, we also explore novel factored observation modeling schemes that allow for greater flexibility in combining the tandem and standard features.
Keywords
hidden Markov models; multilayer perceptrons; natural language processing; speech recognition; MLPS; Mandarin language; articulatory feature classification; automatic speech recognition; cross-language condition; matched-language condition; multilayer perceptron; portable phone; tandem features; Automatic speech recognition; Data mining; Feature extraction; Feedforward neural networks; Hidden Markov models; Multilayer perceptrons; Natural languages; Neural networks; Speech recognition; Telephony; Speech recognition; feedforward neural networks; hidden Markov models;
fLanguage
English
Publisher
ieee
Conference_Titel
Automatic Speech Recognition & Understanding, 2007. ASRU. IEEE Workshop on
Conference_Location
Kyoto
Print_ISBN
978-1-4244-1746-9
Electronic_ISBN
978-1-4244-1746-9
Type
conf
DOI
10.1109/ASRU.2007.4430080
Filename
4430080
Link To Document