Title :
Model-aided coding: a new approach to incorporate facial animation into motion-compensated video coding
Author :
Eisert, Peter ; Wiegand, Thomas ; Girod, Bernd
Author_Institution :
Telecommun. Lab., Erlangen-Nurnberg Univ., Germany
fDate :
4/1/2000 12:00:00 AM
Abstract :
We show that traditional waveform coding and 3-D model-based coding are not competing alternatives, but should be combined to support and complement each other. Both approaches are combined such that the generality of waveform coding and the efficiency of 3-D model-based coding are available where needed. The combination is achieved by providing the block-based video coder with a second reference frame for prediction, which is synthesized by the model-based coder. The model-based coder uses a parameterized 3-D head model, specifying the shape and color of a person. We therefore restrict our investigations to typical videotelephony scenarios that show head-and-shoulder scenes. Motion and deformation of the 3-D head model constitute facial expressions which are represented by facial animation parameters (FAPs) based on the MPEG-4 standard. An intensity gradient-based approach that exploits the 3-D model information is used to estimate the FAPs, as well as illumination parameters, that describe changes of the brightness in the scene. Model failures and objects that are not known at the decoder are handled by standard block-based motion-compensated prediction, which is not restricted to a special scene content, but results in lower coding efficiency. A Lagrangian approach is employed to determine the most efficient prediction for each block from either the synthesized model frame or the previous decoded frame. Experiments on five video sequences show that bit rate savings of about 35% are achieved at equal average peak signal-to-noise ratio (PSNR) when comparing the model-aided codec to TMN-10, the state-of-the-art test model of the M.263 standard. This corresponds to a gain of 2-3 dB in PSNR when encoding at the same average bit rate
Keywords :
computer animation; gradient methods; image sequences; motion compensation; prediction theory; video codecs; video coding; videotelephony; 3D head model deformation; 3D head model motion; 3D model-based coding; Lagrangian approach; M.263 standard; MPEG-4 standard; PSNR; TMN-10; bit rate savings; block-based motion-compensated prediction; block-based video coder; brightness; coding efficiency; decoder; experiments; facial animation parameters; facial expressions; head-and-shoulder scenes; illumination parameters; intensity gradient; model-aided codec; model-aided coding; motion-compensated video coding; parameterized 3D head model; peak signal-to-noise ratio; prediction; reference frame; synthesized model frame; test model; video sequences; videotelephony; waveform coding; Bit rate; Decoding; Deformable models; Facial animation; Financial advantage program; Layout; MPEG 4 Standard; PSNR; Predictive models; Shape;
Journal_Title :
Circuits and Systems for Video Technology, IEEE Transactions on