Systematic evaluation of CNN advances on the ImageNet
arXiv:1606.02228 · doi:10.1016/j.cviu.2017.05.007
Abstract
The paper systematically studies the impact of a range of recent advances in CNN architectures and learning methods on the object categorization (ILSVRC) problem. The evalution tests the influence of the following choices of the architecture: non-linearity (ReLU, ELU, maxout, compatibility with batch normalization), pooling variants (stochastic, max, average, mixed), network width, classifier design (convolutional, fully-connected, SPP), image pre-processing, and of learning parameters: learning rate, batch size, cleanliness of the data, etc. The performance gains of the proposed modifications are first tested individually and then in combination. The sum of individual gains is bigger than the observed improvement when all modifications are introduced, but the "deficit" is small suggesting independence of their benefits. We show that the use of 128x128 pixel images is sufficient to make qualitative conclusions about optimal network structure that hold for the full size Caffe and VGG nets. The results are obtained an order of magnitude faster than with the standard 224 pixel images.
Submitted to CVIU Special Issue on Deep Learning. Updated dataset quality experiment
References in corpus (19)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Neural Architecture Search with Reinforcement Learning
- Striving for Simplicity: The All Convolutional Net
- Wide Residual Networks
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
- One weird trick for parallelizing convolutional neural networks
- An Analysis of Deep Neural Network Models for Practical Applications
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- Understanding the Effective Receptive Field in Deep Convolutional Neural Networks
- The Loss Surfaces of Multilayer Networks
- FractalNet: Ultra-Deep Neural Networks without Residuals
- Stochastic Pooling for Regularization of Deep Convolutional Neural Networks
- Learning Activation Functions to Improve Deep Neural Networks
- Practical recommendations for gradient-based training of deep architectures
- No bad local minima: Data independent training error guarantees for multilayer neural networks
- Multi-Bias Non-linear Activation in Deep Neural Networks
- On architectural choices in deep learning: From network structure to gradient convergence and parameter estimation
Cited by in corpus (43)
- Deep learning with convolutional neural networks for EEG decoding and visualization
- Ensemble Adversarial Training: Attacks and Defenses
- A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
- A Comprehensive Analysis of Deep Regression
- Evaluation of Retinal Image Quality Assessment Networks in Different Color-spaces
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- Detecting and interpreting myocardial infarction using fully convolutional neural networks
- Task-Driven Convolutional Recurrent Models of the Visual System
- Review: Deep Learning in Electron Microscopy
- Fine-Grained Object Recognition and Zero-Shot Learning in Remote Sensing Imagery
- Applying deep neural networks to the detection and space parameter estimation of compact binary coalescence with a network of gravitational wave detectors
- Analysis and Optimization of Convolutional Neural Network Architectures
- Machine Learning Techniques for Stellar Light Curve Classification
- Detecting Adversarial Examples by Input Transformations, Defense Perturbations, and Voting
- PolyNet: A Pursuit of Structural Diversity in Very Deep Networks
- Deep Fingerprinting: Undermining Website Fingerprinting Defenses with Deep Learning
- PreVIous: A Methodology for Prediction of Visual Inference Performance on IoT Devices
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource Constraints
- DDCNet: Deep Dilated Convolutional Neural Network for Dense Prediction
- A Survey on Machine Learning Techniques for Auto Labeling of Video, Audio, and Text Data
- Deep Neural Networks for Marine Debris Detection in Sonar Images
- Agriculture-Vision: A Large Aerial Image Database for Agricultural Pattern Analysis
- Batch Layer Normalization, A new normalization layer for CNNs and RNN
- Empirical Upper Bound in Object Detection and More
- Active Learning for Visual Question Answering: An Empirical Study
- Analysis of Filter Size Effect In Deep Learning
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Identifying Light-curve Signals with a Deep Learning Based Object Detection Algorithm. II. A General Light Curve Classification Framework
- Learning to play the Chess Variant Crazyhouse above World Champion Level with Deep Neural Networks and Human Data
- Learning from Experience for Rapid Generation of Local Car Maneuvers
- EcoNAS: Finding Proxies for Economical Neural Architecture Search
- Deep Face Recognition Model Compression via Knowledge Transfer and Distillation
- Improving the HardNet Descriptor
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- Testing the Efficient Network TRaining (ENTR) Hypothesis: initially reducing training image size makes Convolutional Neural Network training for image recognition tasks more efficient
- Piecewise Linear Units Improve Deep Neural Networks
- Multilevel Context Representation for Improving Object Recognition
- Master's Thesis : Deep Learning for Visual Recognition
- A Resizable Mini-batch Gradient Descent based on a Multi-Armed Bandit
- ScaleNet: An Unsupervised Representation Learning Method for Limited Information
- cGANs for Cartoon to Real-life Images
- Empirical Upper Bound, Error Diagnosis and Invariance Analysis of Modern Object Detectors
- MULTIMODAL ANALYSIS: Informed content estimation and audio source separation