Return of the Devil in the Details: Delving Deep into Convolutional Nets
arXiv:1405.3531
Abstract
The latest generation of Convolutional Neural Networks (CNN) have achieved impressive results in challenging benchmarks on image recognition and object detection, significantly raising the interest of the community in these methods. Nevertheless, it is still unclear how different CNN methods compare with each other and with previous state-of-the-art shallow representations such as the Bag-of-Visual-Words and the Improved Fisher Vector. This paper conducts a rigorous evaluation of these new techniques, exploring different deep architectures and comparing them on a common ground, identifying and disclosing important implementation details. We identify several useful properties of CNN-based representations, including the fact that the dimensionality of the CNN output layer can be reduced significantly without having an adverse effect on performance. We also identify aspects of deep and shallow methods that can be successfully shared. In particular, we show that the data augmentation techniques commonly applied to CNN-based methods can also be applied to shallow methods, and result in an analogous performance boost. Source code and models to reproduce the experiments in the paper is made publicly available.
Published in proceedings of BMVC 2014
Cited by in corpus (65)
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
- Transductive Multi-view Zero-Shot Learning
- Modeling and Propagating CNNs in a Tree Structure for Visual Tracking
- What-and-Where to Match: Deep Spatially Multiplicative Integration Networks for Person Re-identification
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- Robustness of classifiers: from adversarial to random noise
- Fast Feature Fool: A data independent approach to universal adversarial perturbations
- Wavelet Convolutional Neural Networks for Texture Classification
- Universal adversarial perturbations
- DiracNets: Training Very Deep Neural Networks Without Skip-Connections
- Deep convolutional filter banks for texture recognition and segmentation
- Feedforward semantic segmentation with zoom-out features
- Delta Networks for Optimized Recurrent Network Computation
- Stacked Quantizers for Compositional Vector Compression
- Towards Automated Melanoma Screening: Exploring Transfer Learning Schemes
- Fisher Kernel for Deep Neural Activations
- Improved Bilinear Pooling with CNNs
- Simple Image Description Generator via a Linear Phrase-Based Approach
- A Discriminative CNN Video Representation for Event Detection
- Convolutional Neural Networks at Constrained Time Cost
- When Unsupervised Domain Adaptation Meets Tensor Representations
- Soft Proposal Networks for Weakly Supervised Object Localization
- What's Mine is Yours: Pretrained CNNs for Limited Training Sonar ATR
- Video-based Human Action Recognition using Deep Learning: A Review
- Deep Learning for Object Saliency Detection and Image Segmentation
- Not Afraid of the Dark: NIR-VIS Face Recognition via Cross-spectral Hallucination and Low-rank Embedding
- HashGAN:Attention-aware Deep Adversarial Hashing for Cross Modal Retrieval
- Privacy-Preserving Deep Inference for Rich User Data on The Cloud
- Automatic Classification of Bright Retinal Lesions via Deep Network Features
- Convolutional Neural Network-Based Image Representation for Visual Loop Closure Detection
- Temporal Dynamic Graph LSTM for Action-driven Video Object Detection
- ScreenAvoider: Protecting Computer Screens from Ubiquitous Cameras
- A Jointly Learned Deep Architecture for Facial Attribute Analysis and Face Detection in the Wild
- Recurrent Filter Learning for Visual Tracking
- Cultural Event Recognition with Visual ConvNets and Temporal Models
- Holistic Interstitial Lung Disease Detection using Deep Convolutional Neural Networks: Multi-label Learning and Unordered Pooling
- Fine-tuning deep CNN models on specific MS COCO categories
- Multi-Modal Music Information Retrieval: Augmenting Audio-Analysis with Visual Computing for Improved Music Video Analysis
- Object-Scene Convolutional Neural Networks for Event Recognition in Images
- Multi-hypothesis contextual modeling for semantic segmentation
- Asymmetric Deep Supervised Hashing
- Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video Classification
- Where to Focus: Deep Attention-based Spatially Recurrent Bilinear Networks for Fine-Grained Visual Recognition
- Visual Summary of Egocentric Photostreams by Representative Keyframes
- How hard can it be? Estimating the difficulty of visual search in an image
- Domain-Size Pooling in Local Descriptors: DSP-SIFT
- Bi-class classification of humpback whale sound units against complex background noise with Deep Convolution Neural Network
- Material Classification in the Wild: Do Synthesized Training Data Generalise Better than Real-World Training Data?
- Exploiting Multi-modal Curriculum in Noisy Web Data for Large-scale Concept Learning
- Optimizing Filter Size in Convolutional Neural Networks for Facial Action Unit Recognition
- Semantic tracking: Single-target tracking with inter-supervised convolutional networks
- Towards CNN Map Compression for camera relocalisation
- Learning Deep Representations for Scene Labeling with Semantic Context Guided Supervision
- Geometric Convolutional Neural Network for Analyzing Surface-Based Neuroimaging Data
- Make Your Bone Great Again : A study on Osteoporosis Classification
- Real-time Online Action Detection Forests using Spatio-temporal Contexts
- Efficient On-the-fly Category Retrieval using ConvNets and GPUs
- Higher-order Pooling of CNN Features via Kernel Linearization for Action Recognition
- Learning Robust Hash Codes for Multiple Instance Image Retrieval
- Deep Collaborative Learning for Visual Recognition
- Dynamic texture and scene classification by transferring deep image features
- Weakly supervised segment annotation via expectation kernel density estimation
- MV-C3D: A Spatial Correlated Multi-View 3D Convolutional Neural Networks
- Gram Regularization for Multi-view 3D Shape Retrieval
- An Analysis of Human-centered Geolocation