Beyond the Informedia digital video library: video and audio analysis for remembering conversations

Author

Hauptmann, Alexander G. ; Lin, Wei-Hao

Author_Institution

Sch. of Comput. Sci., Carnegie Mellon Univ., Pittsburgh, PA, USA

fYear

2001

fDate

2001

Firstpage

296

Lastpage

300

Abstract

The Informedia Project digital video library pioneered the automatic analysis of television broadcast news and its retrieval on demand. Building on that system, we have developed a wearable, personalized Informedia system, which listens to and transcribes the wearer\´s part of a conversation, recognizes the face of the current dialog partner and remembers his/her voice. The next time the system sees the same person\´s face and hears the same voice, it can retrieve the audio from the last conversation, replaying in compressed form the names and major issues that were mentioned. All of this happens unobtrusively, somewhat like an intelligent assistant who whispers to you: "That\´s Bob Jones from Tech Solutions; two weeks ago in London you discussed solar panels". This paper outlines the general system components as well as interface considerations. Initial implementations showed that both face recognition methods and speaker identification technology have serious shortfalls that must be overcome.

Keywords

audio databases; face recognition; information retrieval; mobile computing; natural language interfaces; personal information systems; speaker recognition; video databases; Informedia Project; audio analysis; digital video library; face recognition; interface considerations; personalized Informedia system; speaker identification; video analysis; wearable computing; Automatic speech recognition; Computer science; Digital audio broadcasting; Digital video broadcasting; Face recognition; Global Positioning System; Software libraries; Space technology; Speech recognition; TV broadcasting;

fLanguage

English

Publisher

ieee

Conference_Titel

Automatic Speech Recognition and Understanding, 2001. ASRU '01. IEEE Workshop on

Print_ISBN

0-7803-7343-X

Type

conf

DOI

10.1109/ASRU.2001.1034646

Filename

1034646