• DocumentCode
    5939
  • Title

    3D Convolutional Neural Networks for Human Action Recognition

  • Author

    Shuiwang Ji ; Wei Xu ; Ming Yang ; Kai Yu

  • Author_Institution
    Dept. of Comput. Sci., Old Dominion Univ., Norfolk, VA, USA
  • Volume
    35
  • Issue
    1
  • fYear
    2013
  • fDate
    Jan. 2013
  • Firstpage
    221
  • Lastpage
    231
  • Abstract
    We consider the automated recognition of human actions in surveillance videos. Most current methods build classifiers based on complex handcrafted features computed from the raw inputs. Convolutional neural networks (CNNs) are a type of deep model that can act directly on the raw inputs. However, such models are currently limited to handling 2D inputs. In this paper, we develop a novel 3D CNN model for action recognition. This model extracts features from both the spatial and the temporal dimensions by performing 3D convolutions, thereby capturing the motion information encoded in multiple adjacent frames. The developed model generates multiple channels of information from the input frames, and the final feature representation combines information from all channels. To further boost the performance, we propose regularizing the outputs with high-level features and combining the predictions of a variety of different models. We apply the developed models to recognize human actions in the real-world environment of airport surveillance videos, and they achieve superior performance in comparison to baseline methods.
  • Keywords
    feature extraction; gesture recognition; image classification; image motion analysis; image representation; neural nets; spatiotemporal phenomena; video surveillance; 3D CNN model; 3D convolutional neural networks; airport surveillance videos; automated human action recognition; baseline methods; complex handcrafted features; deep model; feature representation; high-level features; motion information encoding; spatial dimensions; temporal dimensions; Computational modeling; Computer architecture; Feature extraction; Kernel; Solid modeling; Three dimensional displays; Videos; 3D convolution; Deep learning; action recognition; convolutional neural networks; model combination; Algorithms; Decision Support Techniques; Image Interpretation, Computer-Assisted; Imaging, Three-Dimensional; Movement; Neural Networks (Computer); Pattern Recognition, Automated; Subtraction Technique;
  • fLanguage
    English
  • Journal_Title
    Pattern Analysis and Machine Intelligence, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    0162-8828
  • Type

    jour

  • DOI
    10.1109/TPAMI.2012.59
  • Filename
    6165309