DocumentCode :
3485413
Title :
From Modern Standard Arabic to Levantine ASR: Leveraging GALE for dialects
Author :
Soltau, Hagen ; Mangu, Lidia ; Biadsy, Fadi
Author_Institution :
IBM T. J. Watson Res. Center, Yorktown Heights, NY, USA
fYear :
2011
fDate :
11-15 Dec. 2011
Firstpage :
266
Lastpage :
271
Abstract :
We report a series of experiments about how we can progress from Modern Standard Arabic (MSA) to Levantine ASR, in the context of the GALE DARPA program. While our GALE models achieved very low error rates, we still see error rates twice as high when decoding dialectal data. In this paper, we make use of a state-of-the-art Arabic dialect recognition system to automatically identify Levantine and MSA subsets in mixed speech of a variety of dialects including MSA. Training separate models on these subsets, we show a significant reduction in word error rate over using the entire data set to train one system for both dialects. During decoding, we use a tree array structure to mix Levantine and MSA models automatically using the posterior probabilities of the dialect classifier as soft weights. This technique allows us to mix these models without sacrificing performance for either varieties. Furthermore, using the initial acoustic-based dialect recognition system´s output, we show that we can bootstrap a text-based dialect classifier and use it to identify relevant text data for building Levantine language models. Moreover, we compare different vowelization approaches when transitioning from MSA to Levantine models.
Keywords :
speech recognition; GALE DARPA program; MSA models; acoustic-based arabic dialect recognition system; building Levantine language; levantine ASR; modern standard arabic; posterior probabilities; text-based dialect classifier; training separate models; vowelization approaches; Acoustics; Adaptation models; Data models; Decoding; Error analysis; Training; Training data;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Automatic Speech Recognition and Understanding (ASRU), 2011 IEEE Workshop on
Conference_Location :
Waikoloa, HI
Print_ISBN :
978-1-4673-0365-1
Electronic_ISBN :
978-1-4673-0366-8
Type :
conf
DOI :
10.1109/ASRU.2011.6163942
Filename :
6163942
Link To Document :
بازگشت