Title :
Video content extraction and representation using a joint audio and video processing
Author :
Saraceno, Caterina
Author_Institution :
PRIP Inst. for Autom., Vienna Univ. of Technol., Austria
Abstract :
Computer technology allows for large collections of digital archived material. At the same time, the increasing availability of potentially interesting data makes difficult the retrieval of desired information. Currently, access to such information is limited to textual queries or characteristics such as color or texture. The demand for new solutions allowing common users to easily access, store and retrieve relevant audio-visual information is becoming urgent. One possible solution to this problem is to hierarchically organize the audio-visual data so as to create a nested indexing structure which provides efficient access to relevant information at each level of the hierarchy. This work presents an automatic methodology to extract and hierarchically represent the semantics of the contents, based on a joint audio and visual analysis. Descriptions on each media (audio, video) are used to recognize higher level of meaningful structures, such as specific types of scenes, or, at the highest level, correlations beyond the temporal organization of information, allowing it to reflect classes of visual or audio or audio-visual types. Once a hierarchy is extracted from the data analysis, a nested indexing structure can be created to access relevant information at a specific level of detail, according to the user requirements
Keywords :
audio signal processing; content-based retrieval; feature extraction; image representation; video databases; video signal processing; audio analysis; audio processing; audio-visual information retrieval; automatic method; color; computer technology; correlations; data analysis; digital archived material; information access; nested indexing structure; textual queries; texture; video content extraction; video content representation; video processing; visual analysis; Automation; Cameras; Data analysis; Data mining; Databases; Indexing; Information retrieval; Layout; Shape; Video sequences;
Conference_Titel :
Acoustics, Speech, and Signal Processing, 1999. Proceedings., 1999 IEEE International Conference on
Conference_Location :
Phoenix, AZ
Print_ISBN :
0-7803-5041-3
DOI :
10.1109/ICASSP.1999.757480