Very Deep Convolutional Networks for Large-Scale Image Recognition
arXiv:1409.1556
Abstract
In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers. These findings were the basis of our ImageNet Challenge 2014 submission, where our team secured the first and the second places in the localisation and classification tracks respectively. We also show that our representations generalise well to other datasets, where they achieve state-of-the-art results. We have made our two best-performing ConvNet models publicly available to facilitate further research on the use of deep visual representations in computer vision.
References in corpus (6)
- Going Deeper with Convolutions
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Deep convolutional filter banks for texture recognition and segmentation
- Material Recognition in the Wild with the Materials in Context Database
- Actions and Attributes from Wholes and Parts
Cited by in corpus (62)
- Striving for Simplicity: The All Convolutional Net
- FitNets: Hints for Thin Deep Nets
- Learning Face Representation from Scratch
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Compressing Deep Convolutional Networks using Vector Quantization
- DeepID3: Face Recognition with Very Deep Neural Networks
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- An Empirical Evaluation of Deep Learning on Highway Driving
- Massively Parallel Methods for Deep Reinforcement Learning
- Explain Images with Multimodal Recurrent Neural Networks
- Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition
- Deep Image: Scaling up Image Recognition
- Fully Convolutional Multi-Class Multiple Instance Learning
- Fully Connected Deep Structured Networks
- Deep Gaze I: Boosting Saliency Prediction with Feature Maps Trained on ImageNet
- Learning to Compare Image Patches via Convolutional Neural Networks
- Spatially-sparse convolutional neural networks
- Training Deeper Convolutional Networks with Deep Supervision
- Exploring Nearest Neighbor Approaches for Image Captioning
- End-to-End Photo-Sketch Generation via Fully Convolutional Representation Learning
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Evaluating Two-Stream CNN for Video Classification
- Scale-Invariant Convolutional Neural Networks
- Move Evaluation in Go Using Deep Convolutional Neural Networks
- Exploiting Local Features from Deep Networks for Image Retrieval
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- Visual Madlibs: Fill in the blank Image Generation and Question Answering
- DeepID-Net: Deformable Deep Convolutional Neural Networks for Object Detection
- Exploiting Image-trained CNN Architectures for Unconstrained Video Classification
- segDeepM: Exploiting Segmentation and Context in Deep Neural Networks for Object Detection
- Deep convolutional filter banks for texture recognition and segmentation
- Visual Causal Feature Learning
- Interleaved Text/Image Deep Mining on a Large-Scale Radiology Database for Automated Image Interpretation
- When Face Recognition Meets with Deep Learning: an Evaluation of Convolutional Neural Networks for Face Recognition
- Untangling Local and Global Deformations in Deep Convolutional Networks for Image Classification and Sliding Window Detection
- Compressing Convolutional Neural Networks
- Jointly Modeling Embedding and Translation to Bridge Video and Language
- A Discriminative CNN Video Representation for Event Detection
- Convolutional Neural Networks at Constrained Time Cost
- Viewpoints and Keypoints
- Joint Object and Part Segmentation using Deep Learned Potentials
- Modeling Spatial-Temporal Clues in a Hybrid Deep Learning Framework for Video Classification
- Beyond Frontal Faces: Improving Person Recognition Using Multiple Cues
- Transfer Learning for Video Recognition with Scarce Training Data for Deep Convolutional Neural Network
- Beyond Gaussian Pyramid: Multi-skip Feature Stacking for Action Recognition
- Hypercolumns for Object Segmentation and Fine-grained Localization
- Material Recognition in the Wild with the Materials in Context Database
- Boosting Convolutional Features for Robust Object Proposals
- Deep Learning for Object Saliency Detection and Image Segmentation
- Half-CNN: A General Framework for Whole-Image Regression
- Convolutional Neural Network-Based Image Representation for Visual Loop Closure Detection
- The Application of Two-level Attention Models in Deep Convolutional Neural Network for Fine-grained Image Classification
- Actions and Attributes from Wholes and Parts
- Cultural Event Recognition with Visual ConvNets and Temporal Models
- Learning language through pictures
- Object-centric Sampling for Fine-grained Image Classification
- Object-Scene Convolutional Neural Networks for Event Recognition in Images
- Measuring and Understanding Sensory Representations within Deep Networks Using a Numerical Optimization Framework
- Co-Regularized Deep Representations for Video Summarization
- Learning Temporal Embeddings for Complex Video Analysis
- Multi-scale recognition with DAG-CNNs
- Compression Artifacts Reduction by a Deep Convolutional Network