Return of the Devil in the Details: Delving Deep into Convolutional Nets
arXiv:1405.3531
Abstract
The latest generation of Convolutional Neural Networks (CNN) have achieved impressive results in challenging benchmarks on image recognition and object detection, significantly raising the interest of the community in these methods. Nevertheless, it is still unclear how different CNN methods compare with each other and with previous state-of-the-art shallow representations such as the Bag-of-Visual-Words and the Improved Fisher Vector. This paper conducts a rigorous evaluation of these new techniques, exploring different deep architectures and comparing them on a common ground, identifying and disclosing important implementation details. We identify several useful properties of CNN-based representations, including the fact that the dimensionality of the CNN output layer can be reduced significantly without having an adverse effect on performance. We also identify aspects of deep and shallow methods that can be successfully shared. In particular, we show that the data augmentation techniques commonly applied to CNN-based methods can also be applied to shallow methods, and result in an analogous performance boost. Source code and models to reproduce the experiments in the paper is made publicly available.
Published in proceedings of BMVC 2014
Cited by in corpus (66)
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
- Transductive Multi-view Zero-Shot Learning
- Modeling and Propagating CNNs in a Tree Structure for Visual Tracking
- What-and-Where to Match: Deep Spatially Multiplicative Integration Networks for Person Re-identification
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- Robustness of classifiers: from adversarial to random noise
- Fast Feature Fool: A data independent approach to universal adversarial perturbations
- Wavelet Convolutional Neural Networks for Texture Classification
- Universal adversarial perturbations
- DiracNets: Training Very Deep Neural Networks Without Skip-Connections
- Automatic identification of fossils and abiotic grains during carbonate microfacies analysis using deep convolutional neural networks
- Deep convolutional filter banks for texture recognition and segmentation
- Feedforward semantic segmentation with zoom-out features
- Delta Networks for Optimized Recurrent Network Computation
- Stacked Quantizers for Compositional Vector Compression
- Towards Automated Melanoma Screening: Exploring Transfer Learning Schemes
- Fisher Kernel for Deep Neural Activations
- Improved Bilinear Pooling with CNNs
- A Discriminative CNN Video Representation for Event Detection
- Simple Image Description Generator via a Linear Phrase-Based Approach
- Convolutional Neural Networks at Constrained Time Cost
- When Unsupervised Domain Adaptation Meets Tensor Representations
- What's Mine is Yours: Pretrained CNNs for Limited Training Sonar ATR
- Video-based Human Action Recognition using Deep Learning: A Review
- Soft Proposal Networks for Weakly Supervised Object Localization
- Deep Learning for Object Saliency Detection and Image Segmentation
- Not Afraid of the Dark: NIR-VIS Face Recognition via Cross-spectral Hallucination and Low-rank Embedding
- HashGAN:Attention-aware Deep Adversarial Hashing for Cross Modal Retrieval
- Privacy-Preserving Deep Inference for Rich User Data on The Cloud
- Automatic Classification of Bright Retinal Lesions via Deep Network Features
- Convolutional Neural Network-Based Image Representation for Visual Loop Closure Detection
- Temporal Dynamic Graph LSTM for Action-driven Video Object Detection
- ScreenAvoider: Protecting Computer Screens from Ubiquitous Cameras
- A Jointly Learned Deep Architecture for Facial Attribute Analysis and Face Detection in the Wild
- Recurrent Filter Learning for Visual Tracking
- Cultural Event Recognition with Visual ConvNets and Temporal Models
- Holistic Interstitial Lung Disease Detection using Deep Convolutional Neural Networks: Multi-label Learning and Unordered Pooling
- Fine-tuning deep CNN models on specific MS COCO categories
- Multi-Modal Music Information Retrieval: Augmenting Audio-Analysis with Visual Computing for Improved Music Video Analysis
- Asymmetric Deep Supervised Hashing
- Multi-hypothesis contextual modeling for semantic segmentation
- Object-Scene Convolutional Neural Networks for Event Recognition in Images
- Visual Summary of Egocentric Photostreams by Representative Keyframes
- Domain-Size Pooling in Local Descriptors: DSP-SIFT
- Where to Focus: Deep Attention-based Spatially Recurrent Bilinear Networks for Fine-Grained Visual Recognition
- Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video Classification
- How hard can it be? Estimating the difficulty of visual search in an image
- Bi-class classification of humpback whale sound units against complex background noise with Deep Convolution Neural Network
- Exploiting Multi-modal Curriculum in Noisy Web Data for Large-scale Concept Learning
- Optimizing Filter Size in Convolutional Neural Networks for Facial Action Unit Recognition
- Material Classification in the Wild: Do Synthesized Training Data Generalise Better than Real-World Training Data?
- Towards CNN Map Compression for camera relocalisation
- Learning Deep Representations for Scene Labeling with Semantic Context Guided Supervision
- Geometric Convolutional Neural Network for Analyzing Surface-Based Neuroimaging Data
- Semantic tracking: Single-target tracking with inter-supervised convolutional networks
- Make Your Bone Great Again : A study on Osteoporosis Classification
- Real-time Online Action Detection Forests using Spatio-temporal Contexts
- Higher-order Pooling of CNN Features via Kernel Linearization for Action Recognition
- Efficient On-the-fly Category Retrieval using ConvNets and GPUs
- Dynamic texture and scene classification by transferring deep image features
- Deep Collaborative Learning for Visual Recognition
- Learning Robust Hash Codes for Multiple Instance Image Retrieval
- Weakly supervised segment annotation via expectation kernel density estimation
- Gram Regularization for Multi-view 3D Shape Retrieval
- MV-C3D: A Spatial Correlated Multi-View 3D Convolutional Neural Networks
- An Analysis of Human-centered Geolocation