Do Deep Convolutional Nets Really Need to be Deep and Convolutional?
arXiv:1603.05691
Abstract
Yes, they do. This paper provides the first empirical demonstration that deep convolutional models really need to be both deep and convolutional, even when trained with methods such as distillation that allow small or shallow models of high accuracy to be trained. Although previous research showed that shallow feed-forward nets sometimes can learn the complex functions previously learned by deep nets while using the same number of parameters as the deep models they mimic, in this paper we demonstrate that the same methods cannot be used to train accurate models on CIFAR-10 unless the student models contain multiple layers of convolution. Although the student models do not have to be as deep as the teacher model they mimic, the students need multiple convolutional layers to learn functions of comparable accuracy as the deep convolutional teacher.
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Practical Bayesian Optimization of Machine Learning Algorithms
- Deep Residual Learning for Image Recognition
- Understanding deep learning requires rethinking generalization
- Scalable Bayesian Optimization Using Deep Neural Networks
- Transferring Knowledge from a RNN to a DNN
- Big Neural Networks Waste Capacity
- Convolutional Rectifier Networks as Generalized Tensor Decompositions
Cited by in corpus (44)
- Knowledge Distillation: A Survey
- FractalNet: Ultra-Deep Neural Networks without Residuals
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Born Again Neural Networks
- N2N Learning: Network to Network Compression via Policy Gradient Reinforcement Learning
- Light-Weight RefineNet for Real-Time Semantic Segmentation
- Deep Model Compression: Distilling Knowledge from Noisy Teachers
- ISeeU: Visually interpretable deep learning for mortality prediction inside the ICU
- Compressing GANs using Knowledge Distillation
- Efficient Processing of Deep Neural Networks: A Tutorial and Survey
- Learning Approximate Inference Networks for Structured Prediction
- Knowledge Distillation via Route Constrained Optimization
- Data-driven emergence of convolutional structure in neural networks
- FEED: Feature-level Ensemble for Knowledge Distillation
- SAGE: A Split-Architecture Methodology for Efficient End-to-End Autonomous Vehicle Control
- Active Long Term Memory Networks
- Universal Approximation Power of Deep Residual Neural Networks via Nonlinear Control Theory
- Deep Architectures for Modulation Recognition
- Model Distillation with Knowledge Transfer from Face Classification to Alignment and Verification
- Structured Knowledge Distillation for Dense Prediction
- Knowledge Distillation for End-to-End Person Search
- Comparing Kullback-Leibler Divergence and Mean Squared Error Loss in Knowledge Distillation
- Fast, Accurate, and Simple Models for Tabular Data via Augmented Distillation
- Improving the Interpretability of Deep Neural Networks with Knowledge Distillation
- Improving Regression Performance with Distributional Losses
- Noisy Self-Knowledge Distillation for Text Summarization
- Ensemble Knowledge Distillation for Learning Improved and Efficient Networks
- Multi-Attention Based Ultra Lightweight Image Super-Resolution
- A Fast Deep Learning Model for Textual Relevance in Biomedical Information Retrieval
- MS-KD: Multi-Organ Segmentation with Multiple Binary-Labeled Datasets
- Do Normalization Layers in a Deep ConvNet Really Need to Be Distinct?
- Hierarchical Residual Attention Network for Single Image Super-Resolution
- ERNIE-Tiny : A Progressive Distillation Framework for Pretrained Transformer Compression
- Embedded Knowledge Distillation in Depth-Level Dynamic Neural Network
- Multilayer Perceptron Algebra
- Correlation Congruence for Knowledge Distillation
- Deep Hough-Transform Line Priors
- Towards an Understanding of Neural Networks in Natural-Image Spaces
- Are wider nets better given the same number of parameters?
- Distilling Pixel-Wise Feature Similarities for Semantic Segmentation
- Efficient Action Recognition Using Confidence Distillation
- Programming with Neural Surrogates of Programs
- Frugal Machine Learning
- Learning Energy-Based Approximate Inference Networks for Structured Applications in NLP