Lets keep it simple, Using simple architectures to outperform deeper and more complex architectures
arXiv:1608.06037
Abstract
Major winning Convolutional Neural Networks (CNNs), such as AlexNet, VGGNet, ResNet, GoogleNet, include tens to hundreds of millions of parameters, which impose considerable computation and memory overhead. This limits their practical use for training, optimization and memory efficiency. On the contrary, light-weight architectures, being proposed to address this issue, mainly suffer from low accuracy. These inefficiencies mostly stem from following an ad hoc procedure. We propose a simple architecture, called SimpleNet, based on a set of designing principles, with which we empirically show, a well-crafted yet simple and reasonably deep architecture can perform on par with deeper and more complex architectures. SimpleNet provides a good tradeoff between the computation/memory efficiency and the accuracy. Our simple 13-layer architecture outperforms most of the deeper and complex architectures to date such as VGGNet, ResNet, and GoogleNet on several well-known benchmarks while having 2 to 25 times fewer number of parameters and operations. This makes it very handy for embedded systems or systems with computational and memory limitations. We achieved state-of-the-art result on CIFAR10 outperforming several heavier architectures, near state of the art on MNIST and competitive results on CIFAR100 and SVHN. We also outperformed the much larger and deeper architectures such as VGGNet and popular variants of ResNets among others on the ImageNet dataset. Models are made available at: https://github.com/Coderx7/SimpleNet
Added the long-overdue ImageNet results and updated the missed cifar10/100 results from 2018
References in corpus (15)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Improving neural networks by preventing co-adaptation of feature detectors
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Striving for Simplicity: The All Convolutional Net
- FitNets: Hints for Thin Deep Nets
- Densely Connected Convolutional Networks
- Wide Residual Networks
- Understanding Neural Networks Through Deep Visualization
- Stochastic Pooling for Regularization of Deep Convolutional Neural Networks
- Scalable Bayesian Optimization Using Deep Neural Networks
- Deep Image: Scaling up Image Recognition
- Spatially-sparse convolutional neural networks
- Deep Networks with Internal Selective Attention through Feedback Connections
- APAC: Augmented PAttern Classification with Neural Networks
Cited by in corpus (14)
- Pruning by Explaining: A Novel Criterion for Deep Neural Network Pruning
- A framework for large-scale mapping of human settlement extent from Sentinel-2 images via fully convolutional neural networks
- Similarity-based Label Inference Attack against Training and Inference of Split Learning
- Bit Error Robustness for Energy-Efficient DNN Accelerators
- A Simple Dynamic Learning Rate Tuning Algorithm For Automated Training of DNNs
- TanhSoft -- a family of activation functions combining Tanh and Softplus
- Machine Learning Automation Toolbox (MLaut)
- Multi-level Feature Fusion-based CNN for Local Climate Zone Classification from Sentinel-2 Images: Benchmark Results on the So2Sat LCZ42 Dataset
- Machine-learning enables Image Reconstruction and Classification in a "see-through" camera
- Implicitly Maximizing Margins with the Hinge Loss
- Bridging the Band Gap: What Device Physicists Need to Know About Machine Learning
- CAR -- Cityscapes Attributes Recognition A Multi-category Attributes Dataset for Autonomous Vehicles
- Genealogical Population-Based Training for Hyperparameter Optimization
- EIS -- a family of activation functions combining Exponential, ISRU, and Softplus