Moonshine: Distilling with Cheap Convolutions
arXiv:1711.02613
Abstract
Many engineers wish to deploy modern neural networks in memory-limited settings; but the development of flexible methods for reducing memory use is in its infancy, and there is little knowledge of the resulting cost-benefit. We propose structural model distillation for memory reduction using a strategy that produces a student architecture that is a simple transformation of the teacher architecture: no redesign is needed, and the same hyperparameters can be used. Using attention transfer, we provide Pareto curves/tables for distillation of residual networks with four benchmark datasets, indicating the memory versus accuracy payoff. We show that substantial memory savings are possible with very little loss of accuracy, and confirm that distillation provides student network performance that is better than training that student architecture directly on data.
32nd Conference on Neural Information Processing Systems (NeurIPS 2018)
Cited by in corpus (39)
- Knowledge Distillation: A Survey
- Edge Intelligence: Architectures, Challenges, and Applications
- Applications and Techniques for Fast Machine Learning in Science
- EmBench: Quantifying Performance Variations of Deep Neural Networks across Modern Commodity Devices
- MeliusNet: Can Binary Neural Networks Achieve MobileNet-level Accuracy?
- ResKD: Residual-Guided Knowledge Distillation
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Relational Knowledge Distillation
- Synergic Adversarial Label Learning for Grading Retinal Diseases via Knowledge Distillation and Multi-task Learning
- Zero-shot Knowledge Transfer via Adversarial Belief Matching
- Large-Scale Generative Data-Free Distillation
- Group Sparsity: The Hinge Between Filter Pruning and Decomposition for Network Compression
- Knapsack Pruning with Inner Distillation
- Emotion Recognition in Speech using Cross-Modal Transfer in the Wild
- BlockSwap: Fisher-guided Block Substitution for Network Compression on a Budget
- Search for Better Students to Learn Distilled Knowledge
- BoolNet: Minimizing The Energy Consumption of Binary Neural Networks
- Automated Design Space Exploration for optimised Deployment of DNN on Arm Cortex-A CPUs
- CUP: Cluster Pruning for Compressing Deep Neural Networks
- Class-Distribution-Aware Calibration for Long-Tailed Visual Recognition
- Interactive Knowledge Distillation
- Characterising Across-Stack Optimisations for Deep Convolutional Neural Networks
- Anti-Distillation: Improving reproducibility of deep networks
- Weight Pruning via Adaptive Sparsity Loss
- Distilling with Performance Enhanced Students
- Separable Layers Enable Structured Efficient Linear Substitutions
- Training convolutional neural networks with cheap convolutions and online distillation
- Student Network Learning via Evolutionary Knowledge Distillation
- On the Demystification of Knowledge Distillation: A Residual Network Perspective
- Cross-Modal Knowledge Distillation Method for Automatic Cued Speech Recognition
- On the Orthogonality of Knowledge Distillation with Other Techniques: From an Ensemble Perspective
- Adversarial-Based Knowledge Distillation for Multi-Model Ensemble and Noisy Data Refinement
- Learning from a Lightweight Teacher for Efficient Knowledge Distillation
- Substitute Teacher Networks: Learning with Almost No Supervision
- Dilated DenseNets for Relational Reasoning
- Introspective Learning by Distilling Knowledge from Online Self-explanation
- Training on the Edge: The why and the how
- AsymmNet: Towards ultralight convolution neural networks using asymmetrical bottlenecks
- Neighbourhood Distillation: On the benefits of non end-to-end distillation