Title :
Design and collection of acoustic sound data for hands-free speech recognition and sound scene understanding
Author :
Nakamura, Satoshi ; Hiyane, Kazuo ; Asano, Futoshi ; Kaneda, Yutaka ; Yamada, Takeshi ; Nishiura, Takanobu ; Kobayashi, Tetsunori ; Ise, Shiro ; Saruwatari, Hiroshi
Author_Institution :
ATR Spoken Language Translation Labs., Kyoto, Japan
Abstract :
The sound data for open evaluation is necessary for studies such as sound source localization, sound retrieval, sound recognition and hands-free speech recognition in real acoustic environments. This paper reports on our project for acoustic data collection. There are many kinds of sound scenes in real environments. The sound scene is specified by sound sources and room acoustics. The number of combinations of the sound sources, source positions and rooms is huge in real acoustic environments. We assumed that the sound in the environments can be simulated by convolution of the isolated sound sources and impulse responses. As an isolated sound source, hundred kinds of environment sounds and speech sounds are collected. The impulse responses are collected in various acoustic environments. Additionally we collected sounds from a moving source. In this paper, progress of our sound scene database collection project and application to environment sound recognition and hands-free speech recognition are described.
Keywords :
acoustic signal processing; architectural acoustics; database management systems; direction-of-arrival estimation; hidden Markov models; speech recognition; transient response; acoustic data collection; convolution; hands-free speech recognition; impulse responses; moving source; open evaluation; real acoustic environments; room acoustics; sound data; sound recognition; sound retrieval; sound scene database collection; sound scenes; sound source localization; source positions; speech sounds; Acoustic noise; Humans; Image databases; Layout; Loudspeakers; Microphones; Oral communication; Reverberation; Speech recognition; Working environment noise;
Conference_Titel :
Multimedia and Expo, 2002. ICME '02. Proceedings. 2002 IEEE International Conference on
Print_ISBN :
0-7803-7304-9
DOI :
10.1109/ICME.2002.1035537