A unified framework for filterbank and time-frequency basis vectors in ASR frontends

Author

Xiaoyu Liu ; Zahorian, Stephen A.

Author_Institution

Dept. of Electr. & Comput. Eng., Binghamton Univ., Binghamton, NY, USA

fYear

2015

fDate

19-24 April 2015

Firstpage

4659

Lastpage

4663

Abstract

For many years, filterbank have been widely used as one step of frontend feature extraction for Automatic Speech Recognition (ASR). In this paper, we propose a unified framework for ASR frontends, by first moving the nonlinear amplitude scaling, and then combining the filterbank weights with the cosine basis vectors. As part of this framework, we also show that the delta terms used to encode feature dynamics can also be viewed as one realization of a set of “unified” basis vectors over time. With this framework, frontends can be developed, analyzed and evaluated through their basis vectors over frequency and time.

Keywords

channel bank filters; feature extraction; speech recognition; time-frequency analysis; vectors; ASR frontend; automatic speech recognition; cosine basis vectors; feature dynamics encoding; filterbank weight; frontend feature extraction; nonlinear amplitude scaling; time-frequency basis vectors; unified basis vectors; Accuracy; Feature extraction; Filter banks; Mel frequency cepstral coefficient; Speech; Speech recognition; Time-frequency analysis; Filterbank; basis vector; frontend; spectro-temporal; unified;

fLanguage

English

Publisher

ieee

Conference_Titel

Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on

Conference_Location

South Brisbane, QLD

Type

conf

DOI

10.1109/ICASSP.2015.7178854

Filename

7178854