DocumentCode
1426657
Title
Bilingual Experiments on Automatic Recovery of Capitalization and Punctuation of Automatic Speech Transcripts
Author
Batista, Fernando ; Moniz, Helena ; Trancoso, Isabel ; Mamede, Nuno
Author_Institution
Spoken Language Lab. - L2F, INESC-ID, Lisbon, Portugal
Volume
20
Issue
2
fYear
2012
Firstpage
474
Lastpage
485
Abstract
This paper focuses on the tasks of recovering capitalization and punctuation marks from texts without that information, such as spoken transcripts, produced by automatic speech recognition systems. These two practical rich transcription tasks were performed using the same discriminative approach, based on maximum entropy, suitable for on-the-fly usage. Reported experiments were conducted both over Portuguese and English broadcast news data. Both force aligned and automatic transcripts were used, allowing to measure the impact of the speech recognition errors. Capitalized words and named entities are intrinsically related, and are influenced by time variation effects. For that reason, the so-called language dynamics have been addressed for the capitalization task. Language adaptation results indicate, for both languages, that the capitalization performance is affected by the temporal distance between the training and testing data. In what regards the punctuation task, this paper covers the three most frequent punctuation marks: full stop, comma, and question marks. Different methods were explored for improving the baseline results for full stop and comma. The first uses punctuation information extracted from large written corpora. The second applies different levels of linguistic structure, including lexical, prosodic, and speaker related features. The comma detection improved significantly in the first method, thus indicating that it depends more on lexical features. The second method provided even better results, for both languages and both punctuation marks, best results being achieved mainly for full stop. As for question marks, there is a small gain, but differences are not very significant, due to the relatively small number of question marks in the corpora.
Keywords
speech recognition; English broadcast; Portuguese broadcast; automatic speech recognition system; automatic speech transcription; bilingual experiment; capitalization automatic recovery; comma detection; language adaptation; language dynamic; lexical feature; maximum entropy; punctuation information extraction; punctuation mark automatic recovery; speech recognition error impact measurement; Adaptation models; Hidden Markov models; Speech; Speech processing; Training; Training data; Vocabulary; Automatic speech processing; capitalization; language dynamics; natural language processing; punctuation marks; rich transcription;
fLanguage
English
Journal_Title
Audio, Speech, and Language Processing, IEEE Transactions on
Publisher
ieee
ISSN
1558-7916
Type
jour
DOI
10.1109/TASL.2011.2159594
Filename
6135544
Link To Document