• DocumentCode
    1211316
  • Title

    On Acoustic Diversification Front-End for Spoken Language Identification

  • Author

    Sim, Khe Chai ; Li, Haizhou

  • Author_Institution
    Inst. for Infocomm Res., Singapore
  • Volume
    16
  • Issue
    5
  • fYear
    2008
  • fDate
    7/1/2008 12:00:00 AM
  • Firstpage
    1029
  • Lastpage
    1037
  • Abstract
    The parallel phone recognition followed by language model (PPRLM) architecture represents one of the state-of-the-art spoken language identification systems. A PPRLM system comprises multiple parallel subsystems, where each subsystem employs a phone recognizer with a different phone set for a particular language. The phone recognizer extracts phonotactic attributes from the speech input to characterize a language. The multiple parallel subsystems are devised to capture the phonetic diversification available in the speech input. Alternatively, this paper investigates a new approach for building a PPRLM system that aims at improving the acoustic diversification among its parallel subsystems by using multiple acoustic models. These acoustic models are trained on the same speech data with the same phone set but using different model structures and training paradigms. We examine the use of various structured precision (inverse covariance) matrix modeling techniques as well as the maximum likelihood and maximum mutual information training paradigms to produce complementary acoustic models. The results show that acoustic diversification, which requires only one set of phonetically transcribed speech data, yields similar performance improvements compared to phonetic diversification. In addition, further improvements were obtained by combining both diversification factors. The best performing system reported in this paper combined phonetic and acoustic diversifications to achieve EERs of 4.71% and 8.61% on the 2003 and 2005 NIST LRE sets, respectively, compared to 5.77% and 9.94% using phonetic diversification alone.
  • Keywords
    speech recognition; telephony; acoustic diversification front-end; acoustic model; inverse covariance matrix; language model; maximum likelihood information training; maximum mutual information training; parallel phone recognition; phone recognizer; spoken language identification; structured precision matrix; Acoustic modeling; fusion; maximum mutual information (MMI); parallel phone recognition followed by language model (PPRLM); precision matrix modeling; spoken language identification;
  • fLanguage
    English
  • Journal_Title
    Audio, Speech, and Language Processing, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1558-7916
  • Type

    jour

  • DOI
    10.1109/TASL.2008.924150
  • Filename
    4512018