• DocumentCode
    60720
  • Title

    Image Feature Representation of the Subband Power Distribution for Robust Sound Event Classification

  • Author

    Dennis, Jonathan ; Tran, HuyDat ; Chng, Eng Siong

  • Author_Institution
    Inst. for Infocomm Res., Agency for Sci., Technol. & Res., Singapore, Singapore
  • Volume
    21
  • Issue
    2
  • fYear
    2013
  • fDate
    Feb. 2013
  • Firstpage
    367
  • Lastpage
    377
  • Abstract
    The ability to automatically recognize a wide range of sound events in real-world conditions is an important part of applications such as acoustic surveillance and machine hearing. Our approach takes inspiration from both audio and image processing fields, and is based on transforming the sound into a two-dimensional representation, then extracting an image feature for classification. This provided the motivation for our previous work on the spectrogram image feature (SIF). In this paper, we propose a novel method to improve the sound event classification performance in severe mismatched noise conditions. This is based on the subband power distribution (SPD) image - a novel two-dimensional representation that characterizes the spectral power distribution over time in each frequency subband. Here, the high-powered reliable elements of the spectrogram are transformed to a localized region of the SPD, hence can be easily separated from the noise. We then extract an image feature from the SPD, using the same approach as for the SIF, and develop a novel missing feature classification approach based on a nearest neighbor classifier (kNN). We carry out comprehensive experiments on a database of 50 environmental sound classes over a range of challenging noise conditions. The results demonstrate that the SPD-IF is both discriminative over the broad range of sound classes, and robust in severe non-stationary noise.
  • Keywords
    audio signal processing; image classification; image representation; acoustic surveillance; audio processing; image classification; image feature representation; image processing; machine hearing; mismatched noise conditions; nearest neighbor classifier; robust sound event classification; spectrogram image feature; subband power distribution; two-dimensional representation; Feature extraction; Noise; Robustness; Spectrogram; Speech; Time frequency analysis; Sound event classification; missing feature theory; spectrogram; subband power distribution (SPD);
  • fLanguage
    English
  • Journal_Title
    Audio, Speech, and Language Processing, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1558-7916
  • Type

    jour

  • DOI
    10.1109/TASL.2012.2226160
  • Filename
    6338274