DocumentCode :
410006
Title :
A fast and robust speech/music discrimination approach
Author :
Wang, W.Q. ; Gao, Wenzhong ; Ying, D.W.
Author_Institution :
Coll. of Inf. Sci. & Eng., Chinese Acad. of Sci., China
Volume :
3
fYear :
2003
fDate :
15-18 Dec. 2003
Firstpage :
1325
Abstract :
This paper presents a simple and effective approach to discriminate speech and music. First, the proposed modified low energy ratio is extracted from each window-level segment as the only feature. Then the system applied the Bayes MAP classifier to decide the audio class of each segment. Last, based on the fact that the audio types of neighboring segments have very strong relevance, a novel context-based post-decision method is designed to refine the classification results. The proposed method is evaluated on about 5 hours of audio data, which involves clean and noisy speech from various speakers, as well as a wide range of musical content. The experimental results are promising, and a classification accuracy of more than 97% has been achieved despite the low computation complexity of the method.
Keywords :
Bayes methods; audio signal processing; computational complexity; music; signal classification; speech recognition; Bayes MAP classifier; audio classification; audio data; audio segmentation; context-based postdecision method; modified low energy ratio; music discrimination; speech discrimination; Artificial intelligence; Automatic speech recognition; Computers; Data mining; Educational institutions; Frequency modulation; Multiple signal classification; Power engineering and energy; Robustness; Streaming media;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Information, Communications and Signal Processing, 2003 and Fourth Pacific Rim Conference on Multimedia. Proceedings of the 2003 Joint Conference of the Fourth International Conference on
Print_ISBN :
0-7803-8185-8
Type :
conf
DOI :
10.1109/ICICS.2003.1292679
Filename :
1292679
Link To Document :
بازگشت