Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
arXiv:2012.09816
Abstract
We formally study how ensemble of deep learning models can improve test accuracy, and how the superior performance of ensemble can be distilled into a single model using knowledge distillation. We consider the challenging case where the ensemble is simply an average of the outputs of a few independently trained neural networks with the SAME architecture, trained using the SAME algorithm on the SAME data set, and they only differ by the random seeds used in the initialization. We show that ensemble/knowledge distillation in Deep Learning works very differently from traditional learning theory (such as boosting or NTKs, neural tangent kernels). To properly understand them, we develop a theory showing that when data has a structure we refer to as ``multi-view'', then ensemble of independently trained neural networks can provably improve test accuracy, and such superior test accuracy can also be provably distilled into a single model by training a single model to match the output of the ensemble instead of the true label. Our result sheds light on how ensemble works in deep learning in a way that is completely different from traditional theorems, and how the ``dark knowledge'' is hidden in the outputs of the ensemble and can be used in distillation. In the end, we prove that self-distillation can also be viewed as implicitly combining ensemble and knowledge distillation to improve test accuracy.
v2/V3 polishes writing
References in corpus (13)
- Distilling the Knowledge in a Neural Network
- Popular Ensemble Methods: An Empirical Study
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
- Recovery Guarantees for One-hidden-layer Neural Networks
- Learning One-hidden-layer Neural Networks with Landscape Design
- Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
- Diverse Neural Network Learns True Target Functions
- Neural Kernels Without Tangents
- Theoretical properties of the global optimizer of two layer neural network
- Recovery Guarantee of Non-negative Matrix Factorization via Alternating Updates
- Learning Over-Parametrized Two-Layer ReLU Neural Networks beyond NTK
Cited by in corpus (35)
- R-Drop: Regularized Dropout for Neural Networks
- Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey
- Knowledge Distillation in Deep Learning and its Applications
- Multi-Branch Mutual-Distillation Transformer for EEG-Based Seizure Subtype Classification
- Beyond Heart Murmur Detection: Automatic Murmur Grading from Phonocardiogram
- Photoacoustic image synthesis with generative adversarial networks
- Improving Multi-Modal Learning with Uni-Modal Teachers
- Easy Ensemble: Simple Deep Ensemble Learning for Sensor-Based Human Activity Recognition
- FS-BAN: Born-Again Networks for Domain Generalization Few-Shot Classification
- What is Next when Sequential Prediction Meets Implicitly Hard Interaction?
- Assessing Generalization of SGD via Disagreement
- BEBERT: Efficient and Robust Binary Ensemble BERT
- Variant Parallelism: Lightweight Deep Convolutional Models for Distributed Inference on IoT Devices
- Semantic Segmentation in Multiple Adverse Weather Conditions with Domain Knowledge Retention
- Ensemble Transformer for Efficient and Accurate Ranking Tasks: an Application to Question Answering Systems
- Understanding the Generalization of Adam in Learning Neural Networks with Proper Regularization
- Importance Sampling CAMs for Weakly-Supervised Segmentation
- DistillCSE: Distilled Contrastive Learning for Sentence Embeddings
- Teacher's pet: understanding and mitigating biases in distillation
- Students are the Best Teacher: Exit-Ensemble Distillation with Multi-Exits
- Towards Model Agnostic Federated Learning Using Knowledge Distillation
- A Closer Look at Codistillation for Distributed Training
- Multi-label Iterated Learning for Image Classification with Label Ambiguity
- Deep Neural Compression Via Concurrent Pruning and Self-Distillation
- On the One-sided Convergence of Adam-type Algorithms in Non-convex Non-concave Min-max Optimization
- Adaptive Distillation: Aggregating Knowledge from Multiple Paths for Efficient Distillation
- sDREAMER: Self-distilled Mixture-of-Modality-Experts Transformer for Automatic Sleep Staging
- OCHADAI-KYOTO at SemEval-2021 Task 1: Enhancing Model Generalization and Robustness for Lexical Complexity Prediction
- Explaining generalization in deep learning: progress and fundamental limits
- Minimal Sufficient Views: A DNN model making predictions with more evidence has higher accuracy
- Regression Bugs Are In Your Model! Measuring, Reducing and Analyzing Regressions In NLP Model Updates
- Structured Ensembles: an Approach to Reduce the Memory Footprint of Ensemble Methods
- Technical Report of Team GraphMIRAcles in the WikiKG90M-LSC Track of OGB-LSC @ KDD Cup 2021
- Combining Different V1 Brain Model Variants to Improve Robustness to Image Corruptions in CNNs
- Combining Diverse Feature Priors