Deep Model Compression: Distilling Knowledge from Noisy Teachers
arXiv:1610.09650
Abstract
The remarkable successes of deep learning models across various applications have resulted in the design of deeper networks that can solve complex problems. However, the increasing depth of such models also results in a higher storage and runtime complexity, which restricts the deployability of such very deep models on mobile and portable devices, which have limited storage and battery capacity. While many methods have been proposed for deep model compression in recent years, almost all of them have focused on reducing storage complexity. In this work, we extend the teacher-student framework for deep model compression, since it has the potential to address runtime and train time complexity too. We propose a simple methodology to include a noise-based regularizer while training the student from the teacher, which provides a healthy improvement in the performance of the student network. Our experiments on the CIFAR-10, SVHN and MNIST datasets show promising improvement, with the best performance on the CIFAR-10 dataset. We also conduct a comprehensive empirical evaluation of the proposed method under related settings on the CIFAR-10 dataset to show the promise of the proposed approach.
9 pages, 3 figures
References in corpus (3)
Cited by in corpus (19)
- Adaptive Multi-Teacher Multi-level Knowledge Distillation
- Sobolev Training for Neural Networks
- Relational Knowledge Distillation
- Adaptive Neural Network-Based Approximation to Accelerate Eulerian Fluid Simulation
- Model Distillation with Knowledge Transfer from Face Classification to Alignment and Verification
- OpenEI: An Open Framework for Edge Intelligence
- All You Need is a Few Shifts: Designing Efficient Convolutional Neural Networks for Image Classification
- MarginDistillation: distillation for margin-based softmax
- Triplet Distillation for Deep Face Recognition
- ADD: Augmented Disentanglement Distillation Framework for Improving Stock Trend Forecasting
- Correlation Congruence for Knowledge Distillation
- Embracing the Dark Knowledge: Domain Generalization Using Regularized Knowledge Distillation
- KD-Lib: A PyTorch library for Knowledge Distillation, Pruning and Quantization
- ESPN: Extremely Sparse Pruned Networks
- The Pupil Has Become the Master: Teacher-Student Model-Based Word Embedding Distillation with Ensemble Learning
- CHEER: Rich Model Helps Poor Model via Knowledge Infusion
- Implicit Priors for Knowledge Sharing in Bayesian Neural Networks
- Learning Fast Matching Models from Weak Annotations
- A Studious Approach to Semi-Supervised Learning