• DocumentCode
    290265
  • Title

    “Eigenlips” for robust speech recognition

  • Author

    Bregler, Christoph ; Konig, Yochai

  • Author_Institution
    Int. Comput. Sci. Inst., Berkeley, CA, USA
  • Volume
    ii
  • fYear
    1994
  • fDate
    19-22 Apr 1994
  • Abstract
    We improve the performance of a hybrid connectionist speech recognition system by incorporating visual information about the corresponding lip movements. Specifically, we investigate the benefits of adding visual features in the presence of additive noise and crosstalk (cocktail party effect). Our study extends our previous experiments by using a new visual front end, and an alternative architecture for combining the visual and acoustic information. Furthermore, we have extended our recognizer to a multi-speaker, connected letters recognizer. Our results show a significant improvement for the combined architecture (acoustic and visual information) over just the acoustic system in the presence of additive noise and crosstalk
  • Keywords
    acoustic noise; acoustic signal processing; crosstalk; feedforward neural nets; image processing; multilayer perceptrons; speech recognition; speech recognition equipment; Eigenlips; acoustic information; additive noise; cocktail party effect; crosstalk; hybrid connectionist speech recognition; lip movements; multi-speaker connected letters recognizer; robust speech recognition system; visual features; visual front end; visual information; Acoustic distortion; Acoustic noise; Additive noise; Background noise; Computer science; Crosstalk; Deformable models; Image segmentation; Robustness; Speech recognition;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Acoustics, Speech, and Signal Processing, 1994. ICASSP-94., 1994 IEEE International Conference on
  • Conference_Location
    Adelaide, SA
  • ISSN
    1520-6149
  • Print_ISBN
    0-7803-1775-0
  • Type

    conf

  • DOI
    10.1109/ICASSP.1994.389567
  • Filename
    389567