Title :
A fast and robust speech/music discrimination approach
Author :
Wang, W.Q. ; Gao, Wenzhong ; Ying, D.W.
Author_Institution :
Coll. of Inf. Sci. & Eng., Chinese Acad. of Sci., China
Abstract :
This paper presents a simple and effective approach to discriminate speech and music. First, the proposed modified low energy ratio is extracted from each window-level segment as the only feature. Then the system applied the Bayes MAP classifier to decide the audio class of each segment. Last, based on the fact that the audio types of neighboring segments have very strong relevance, a novel context-based post-decision method is designed to refine the classification results. The proposed method is evaluated on about 5 hours of audio data, which involves clean and noisy speech from various speakers, as well as a wide range of musical content. The experimental results are promising, and a classification accuracy of more than 97% has been achieved despite the low computation complexity of the method.
Keywords :
Bayes methods; audio signal processing; computational complexity; music; signal classification; speech recognition; Bayes MAP classifier; audio classification; audio data; audio segmentation; context-based postdecision method; modified low energy ratio; music discrimination; speech discrimination; Artificial intelligence; Automatic speech recognition; Computers; Data mining; Educational institutions; Frequency modulation; Multiple signal classification; Power engineering and energy; Robustness; Streaming media;
Conference_Titel :
Information, Communications and Signal Processing, 2003 and Fourth Pacific Rim Conference on Multimedia. Proceedings of the 2003 Joint Conference of the Fourth International Conference on
Print_ISBN :
0-7803-8185-8
DOI :
10.1109/ICICS.2003.1292679