DocumentCode
1135084
Title
Vocabulary-Independent Indexing of Spontaneous Speech
Author
Yu, Peng ; Chen, Kaijiang ; Ma, Chengyuan ; Seide, Frank
Author_Institution
Microsoft Res. Asia, Beijing, China
Volume
13
Issue
5
fYear
2005
Firstpage
635
Lastpage
643
Abstract
We present a system for vocabulary-independent indexing of spontaneous speech, i.e., neither do we know the vocabulary of a speech recording nor can we predict which query terms for which a user is going to search. The technique can be applied to information retrieval, information extraction, and data mining. Our specific target is search in recorded conversations in the office/information-worker scenario—teleconferences, meetings, presentations, and voice mails. The focus of this paper is on how to index phonetic lattices. We will show that an index should provide expected term frequencies (ETFs) of query terms. Since, at indexing time, it is unknown which phoneme sequences constitute valid query terms, we will introduce an approximation of ETFs of a query\´s phoneme sequence by
-gram phoneme language models, which are estimated on lattices and organized in an inverted index-like structure for fast access. We will discuss ranking, estimation, and integration of phoneme/word hybrid approaches. Compared with an unindexed baseline without approximation, our approximation leads only to a 3.4% relative loss of search accuracy on the Linguistic Data Consortium (LDC) voicemail task. We also propose a two-stage method for locating individual keyword occurences using the above method as a fast match. A 20-times speedup is achieved over unindexed search at under a 2-point accuracy loss. Last, we will briefly introduce a prototype applet based on the above techniques.
-gram phoneme language models, which are estimated on lattices and organized in an inverted index-like structure for fast access. We will discuss ranking, estimation, and integration of phoneme/word hybrid approaches. Compared with an unindexed baseline without approximation, our approximation leads only to a 3.4% relative loss of search accuracy on the Linguistic Data Consortium (LDC) voicemail task. We also propose a two-stage method for locating individual keyword occurences using the above method as a fast match. A 20-times speedup is achieved over unindexed search at under a 2-point accuracy loss. Last, we will briefly introduce a prototype applet based on the above techniques.Keywords
indexing; information retrieval; speech processing; vocabulary; M-gram phoneme language models; data mining; expected term frequencies; information extraction; information retrieval; inverted index-like structure; phoneme sequences; phonetic lattices; speech recording; spontaneous speech; vocabulary-independent indexing; Broadcasting; Data mining; Frequency; Indexing; Information retrieval; Lattices; Prototypes; Speech; Vocabulary; Voice mail; Expected count; expected term frequency; phonetic search; unlimited vocabulary; vocabulary-independent indexing;
fLanguage
English
Journal_Title
Speech and Audio Processing, IEEE Transactions on
Publisher
ieee
ISSN
1063-6676
Type
jour
DOI
10.1109/TSA.2005.851881
Filename
1495444
Link To Document