DocumentCode :
1494012
Title :
Blind Audiovisual Source Separation Based on Sparse Redundant Representations
Author :
Casanovas, Anna Llagostera ; Monaci, Gianluca ; Vandergheynst, Pierre ; Gribonval, Rémi
Author_Institution :
Signal Process. Lab. 2, EPFL, Lausanne, Switzerland
Volume :
12
Issue :
5
fYear :
2010
Firstpage :
358
Lastpage :
371
Abstract :
In this paper, we propose a novel method which is able to detect and separate audiovisual sources present in a scene. Our method exploits the correlation between the video signal captured with a camera and a synchronously recorded one-microphone audio track. In a first stage, audio and video modalities are decomposed into relevant basic structures using redundant representations. Next, synchrony between relevant events in audio and video modalities is quantified. Based on this co-occurrence measure, audiovisual sources are counted and located in the image using a robust clustering algorithm that groups video structures exhibiting strong correlations with the audio. Next periods where each source is active alone are determined and used to build spectral Gaussian mixture models (GMMs) characterizing the sources acoustic behavior. Finally, these models are used to separate the audio signal in periods during which several sources are mixed. The proposed approach has been extensively tested on synthetic and natural sequences composed of speakers and music instruments. Results show that the proposed method is able to successfully detect, localize, separate, and reconstruct present audiovisual sources.
Keywords :
Gaussian processes; audio signal processing; blind source separation; microphones; signal representation; video signal processing; Gaussian mixture models; audio signal; blind audiovisual source separation; microphone audio track; natural sequences; robust clustering algorithm; sparse redundant representations; video signal; Audiovisual processing; Gaussian mixture models; blind source separation; sparse signal representation;
fLanguage :
English
Journal_Title :
Multimedia, IEEE Transactions on
Publisher :
ieee
ISSN :
1520-9210
Type :
jour
DOI :
10.1109/TMM.2010.2050650
Filename :
5466231
Link To Document :
بازگشت