DocumentCode
253923
Title
Learning and Transferring Mid-level Image Representations Using Convolutional Neural Networks
Author
Oquab, Maxime ; Bottou, Leon ; Laptev, Ivan ; Sivic, Josef
Author_Institution
INRIA, Paris, France
fYear
2014
fDate
23-28 June 2014
Firstpage
1717
Lastpage
1724
Abstract
Convolutional neural networks (CNN) have recently shown outstanding image classification performance in the large- scale visual recognition challenge (ILSVRC2012). The success of CNNs is attributed to their ability to learn rich mid-level image representations as opposed to hand-designed low-level features used in other image classification methods. Learning CNNs, however, amounts to estimating millions of parameters and requires a very large number of annotated image samples. This property currently prevents application of CNNs to problems with limited training data. In this work we show how image representations learned with CNNs on large-scale annotated datasets can be efficiently transferred to other visual recognition tasks with limited amount of training data. We design a method to reuse layers trained on the ImageNet dataset to compute mid-level image representation for images in the PASCAL VOC dataset. We show that despite differences in image statistics and tasks in the two datasets, the transferred representation leads to significantly improved results for object and action classification, outperforming the current state of the art on Pascal VOC 2007 and 2012 datasets. We also show promising results for object and action localization.
Keywords
image classification; image representation; neural nets; CNN; ILSVRC2012; ImageNet dataset; PASCAL VOC dataset; Pascal VOC 2007 datasets; Pascal VOC 2012 datasets; action classification; action localization; convolutional neural networks; image classification performance; image statistics; layer reusing; mid-level image representations; object classification; object localization; visual recognition challenge; visual recognition tasks; Computer vision; Image recognition; Image representation; Neural networks; Training; Training data; Visualization;
fLanguage
English
Publisher
ieee
Conference_Titel
Computer Vision and Pattern Recognition (CVPR), 2014 IEEE Conference on
Conference_Location
Columbus, OH
Type
conf
DOI
10.1109/CVPR.2014.222
Filename
6909618
Link To Document