• DocumentCode
    1922986
  • Title

    Vocal emotion recognition in five native languages of Assam using prosodic features in presence of white noise

  • Author

    Kandali, Aditya Bihar ; Routray, Aurobinda ; Basu, Tapan Kumar

  • Author_Institution
    Dept. of Electr. Eng., Jorhat Eng. Coll., Jorhat, India
  • fYear
    2012
  • fDate
    2-3 March 2012
  • Firstpage
    25
  • Lastpage
    25
  • Abstract
    Summary form only given. This work investigates whether vocal emotion expressions of (i) discrete emotion be distinguished from `no-emotion´ (i.e. neutral), (ii) one discrete emotion be distinguished from another, (iii) surprise, which is actually a cognitive component that could be present with any emotion, be also recognized as distinct emotion, (iv) discrete emotion be recognized cross-lingually. This study will enable us to get more information regarding nature and function of emotion. Furthermore, this work will help in developing a generalized vocal emotion recognition system, which will increase the efficiency of human-machine interaction systems. In this work, an emotional speech database consisting of short sentences of six full-blown basic emotions and neutral is created with 140 simulated utterances per speaker of five native languages of Assam. This database is validated by a Listening Test. The Gaussian Mixture Model (GMM) is used as classifier. The performance of the prosodic feature set is computed at sampling frequency of 8.1 kHz from the utterances with and without additive white noise of 5 db and 0 db Signal-to-Noise Ratios (SNRs) under matched noise training and testing condition.
  • Keywords
    Gaussian processes; emotion recognition; natural language processing; signal classification; signal sampling; speech recognition; white noise; Assam; Gaussian mixture model; SNR; additive white noise; classifier; cognitive component; cross-lingual emotion recognition; discrete emotion; distinct emotion recognition; emotional speech database; frequency 8.1 kHz; generalized vocal emotion recognition system; human-machine interaction systems; listening test; matched noise training; native languages; noise figure 0 dB; noise figure 5 dB; prosodic feature set performance; prosodic features; sampling frequency; signal-to-noise ratios; simulated utterances; surprise; testing condition; vocal emotion expressions; Full-blown Basic Emotion; GMM; Prosodic features; Vocal Emotion;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Computational Intelligence and Signal Processing (CISP), 2012 2nd National Conference on
  • Conference_Location
    Guwahati, Assam
  • Print_ISBN
    978-1-4577-0719-3
  • Type

    conf

  • DOI
    10.1109/NCCISP.2012.6189673
  • Filename
    6189673