Knowledge Distillation: A Survey
arXiv:2006.05525 · doi:10.1007/s11263-021-01453-z
Abstract
In recent years, deep neural networks have been successful in both industry and academia, especially for computer vision tasks. The great success of deep learning is mainly due to its scalability to encode large-scale data and to maneuver billions of model parameters. However, it is a challenge to deploy these cumbersome deep models on devices with limited resources, e.g., mobile phones and embedded devices, not only because of the high computational complexity but also the large storage requirements. To this end, a variety of model compression and acceleration techniques have been developed. As a representative type of model compression and acceleration, knowledge distillation effectively learns a small student model from a large teacher model. It has received rapid increasing attention from the community. This paper provides a comprehensive survey of knowledge distillation from the perspectives of knowledge categories, training schemes, teacher-student architecture, distillation algorithms, performance comparison and applications. Furthermore, challenges in knowledge distillation are briefly reviewed and comments on future research are discussed and forwarded.
It has been accepted for publication in International Journal of Computer Vision (2021)
References in corpus (96)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- FitNets: Hints for Thin Deep Nets
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- Ensemble Distillation for Robust Model Fusion in Federated Learning
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- Parallel WaveNet: Fast High-Fidelity Speech Synthesis
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
- Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients
- Adaptive Multi-Teacher Multi-level Knowledge Distillation
- Data-Free Knowledge Distillation for Deep Neural Networks
- Group Knowledge Transfer: Federated Learning of Large CNNs at the Edge
- Semi-supervised Knowledge Transfer for Deep Learning from Private Training Data
- Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
- Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy
- Towards Understanding Knowledge Distillation
- Multilingual Neural Machine Translation with Knowledge Distillation
- N2N Learning: Network to Network Compression via Policy Gradient Reinforcement Learning
- Conditional Teacher-Student Learning
- Deep Model Compression: Distilling Knowledge from Noisy Teachers
- Self-Distillation Amplifies Regularization in Hilbert Space
- Understanding and Improving Knowledge Distillation
- Feature-map-level Online Adversarial Knowledge Distillation
- Comprehensive Attention Self-Distillation for Weakly-Supervised Object Detection
- Compressing GANs using Knowledge Distillation
- FastBERT: a Self-distilling BERT with Adaptive Inference Time
- Transferring Knowledge from a RNN to a DNN
- Self-Distillation as Instance-Specific Label Smoothing
- Relational Knowledge Distillation
- QKD: Quantization-aware Knowledge Distillation
- Online Ensemble Model Compression using Knowledge Distillation
- Knowledge Adaptation: Teaching to Adapt
- Cream of the Crop: Distilling Prioritized Paths For One-Shot Neural Architecture Search
- Object Relational Graph with Teacher-Recommended Learning for Video Captioning
- Graph-based Knowledge Distillation by Multi-head Attention Network
- Knowledge Distillation via Route Constrained Optimization
- Variational Information Distillation for Knowledge Transfer
- Federated Knowledge Distillation
- Doubly Convolutional Neural Networks
- Spatio-Temporal Graph for Video Captioning with Knowledge Distillation
- Regularizing Class-wise Predictions via Self-knowledge Distillation
- TextBrewer: An Open-Source Knowledge Distillation Toolkit for Natural Language Processing
- ShrinkTeaNet: Million-scale Lightweight Face Recognition via Shrinking Teacher-Student Networks
- FEED: Feature-level Ensemble for Knowledge Distillation
- Explaining Sequence-Level Knowledge Distillation as Data-Augmentation for Neural Machine Translation
- Distilling Knowledge from Graph Convolutional Networks
- Knowledge Transfer in Multi-Task Deep Reinforcement Learning for Continuous Control
- In Teacher We Trust: Learning Compressed Models for Pedestrian Detection
- Residual Knowledge Distillation
- Knowledge Adaptation for Efficient Semantic Segmentation
- Cross-Resolution Face Recognition via Prior-Aided Face Hallucination and Residual Knowledge Distillation
- Flexible Dataset Distillation: Learn Labels Instead of Images
- Model Distillation with Knowledge Transfer from Face Classification to Alignment and Verification
- DDFlow: Learning Optical Flow with Unlabeled Data Distillation
- Explaining Knowledge Distillation by Quantifying the Knowledge
- Knowledge distillation via adaptive instance normalization
- Compact Trilinear Interaction for Visual Question Answering
- Search for Better Students to Learn Distilled Knowledge
- Fast, Accurate, and Simple Models for Tabular Data via Augmented Distillation
- Knowledge Squeezed Adversarial Network Compression
- Unpaired Multi-modal Segmentation via Knowledge Distillation
- Student Becoming the Master: Knowledge Amalgamation for Joint Scene Parsing, Depth Estimation, and More
- Dual Policy Distillation
- Progressive Network Grafting for Few-Shot Knowledge Distillation
- Graph Representation Learning via Multi-task Knowledge Distillation
- Improving Neural Architecture Search Image Classifiers via Ensemble Learning
- Inter-Region Affinity Distillation for Road Marking Segmentation
- ALP-KD: Attention-Based Layer Projection for Knowledge Distillation
- Knowledge as Priors: Cross-Modal Knowledge Generalization for Datasets without Superior Knowledge
- Learning Metrics from Teachers: Compact Networks for Image Embedding
- Knowledge Distillation in Wide Neural Networks: Risk Bound, Data Efficiency and Imperfect Teacher
- Creating Something from Nothing: Unsupervised Knowledge Distillation for Cross-Modal Hashing
- Collaborative Teacher-Student Learning via Multiple Knowledge Transfer
- Knowledge Amalgamation from Heterogeneous Networks by Common Feature Learning
- Robust Domain Randomised Reinforcement Learning through Peer-to-Peer Distillation
- Discriminability Distillation in Group Representation Learning
- Diverse Knowledge Distillation for End-to-End Person Search
- Knowledge Transfer via Dense Cross-Layer Mutual-Distillation
- Efficient Video Classification Using Fewer Frames
- Knowledge distillation for optimization of quantized deep neural networks
- Unifying Heterogeneous Classifiers with Distillation
- Two-stage Image Classification Supervised by a Single Teacher Single Student Model
- Peer Collaborative Learning for Online Knowledge Distillation
- Cross-modal knowledge distillation for action recognition
- Knowledge Distillation in Document Retrieval
- Training convolutional neural networks with cheap convolutions and online distillation
- Knowledge Integration Networks for Action Recognition
- Matching Guided Distillation
- Reinforced Multi-Teacher Selection for Knowledge Distillation
- Knowledge Distillation For Recurrent Neural Network Language Modeling With Trust Regularization
- Prime-Aware Adaptive Distillation
- Defocus Blur Detection via Depth Distillation
- Knowledge Representing: Efficient, Sparse Representation of Prior Knowledge for Knowledge Distillation
Cited by in corpus (202)
- Structured Pruning for Deep Convolutional Neural Networks: A survey
- A Comprehensive Survey of Deep Transfer Learning for Anomaly Detection in Industrial Time Series: Methods, Applications, and Directions
- Transformers and Large Language Models for Efficient Intrusion Detection Systems: A Comprehensive Survey
- A Comprehensive Survey of Dataset Distillation
- TrAISformer -- A Transformer Network with Sparse Augmented Data Representation and Cross Entropy Loss for AIS-based Vessel Trajectory Prediction
- Artificial Neural Networks for Photonic Applications: From Algorithms to Implementation
- Deep Long-Tailed Learning: A Survey
- Data Minimization for GDPR Compliance in Machine Learning Models
- Reference-guided Pseudo-Label Generation for Medical Semantic Segmentation
- Learning from models beyond fine-tuning
- On Representation Knowledge Distillation for Graph Neural Networks
- Unified Multi-Modal Image Synthesis for Missing Modality Imputation
- ResKD: Residual-Guided Knowledge Distillation
- MRI-based Alzheimer's disease prediction via distilling the knowledge in multi-modal data
- Reducing Computational Complexity of Neural Networks in Optical Channel Equalization: From Concepts to Implementation
- Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
- Toward Extremely Lightweight Distracted Driver Recognition With Distillation-Based Neural Architecture Search and Knowledge Transfer
- Guided Hybrid Quantization for Object detection in Multimodal Remote Sensing Imagery via One-to-one Self-teaching
- Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
- A Survey on Symbolic Knowledge Distillation of Large Language Models
- Knowledge Distillation approach towards Melanoma Detection
- DnS: Distill-and-Select for Efficient and Accurate Video Indexing and Retrieval
- On-device Training: A First Overview on Existing Systems
- Distilling Object Detectors with Feature Richness
- One-shot Federated Learning without Server-side Training
- Making LLMs Worth Every Penny: Resource-Limited Text Classification in Banking
- Introspective Distillation for Robust Question Answering
- Greening Large Language Models of Code
- Remote Sensing SpatioTemporal Vision-Language Models: A Comprehensive Survey
- Improving Cross-lingual Information Retrieval on Low-Resource Languages via Optimal Transport Distillation
- Trustable Co-label Learning from Multiple Noisy Annotators
- Architectural Vision for Quantum Computing in the Edge-Cloud Continuum
- Low-Resolution Object Recognition with Cross-Resolution Relational Contrastive Distillation
- Distillation Matters: Empowering Sequential Recommenders to Match the Performance of Large Language Model
- Learning Privacy-Preserving Student Networks via Discriminative-Generative Distillation
- Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experiences
- AdaTerm: Adaptive T-Distribution Estimated Robust Moments for Noise-Robust Stochastic Gradient Optimization
- Distilling and Transferring Knowledge via cGAN-generated Samples for Image Classification and Regression
- Edge-aware Feature Aggregation Network for Polyp Segmentation
- Neural-IMLS: Self-supervised Implicit Moving Least-Squares Network for Surface Reconstruction
- Similarity of Neural Network Models: A Survey of Functional and Representational Measures
- Exploring Social Media for Early Detection of Depression in COVID-19 Patients
- An Empirical Study of Challenges in Machine Learning Asset Management
- Computer Vision Model Compression Techniques for Embedded Systems: A Survey
- Theoretical research on generative diffusion models: an overview
- A Dimensionality Reduction Approach for Convolutional Neural Networks
- Towards On-Board Panoptic Segmentation of Multispectral Satellite Images
- Knowledge Distillation in Federated Learning: a Survey on Long Lasting Challenges and New Solutions
- Fly-Swat or Cannon? Cost-Effective Language Model Choice via Meta-Modeling
- Unveiling the frontiers of deep learning: innovations shaping diverse domains
- RankFormer: Listwise Learning-to-Rank Using Listwide Labels
- A Domain-Agnostic Approach for Characterization of Lifelong Learning Systems
- Self-Distribution Binary Neural Networks
- A Comprehensive Survey on Imbalanced Data Learning
- Dynamic Sparse Learning: A Novel Paradigm for Efficient Recommendation
- Synthetic data generation method for data-free knowledge distillation in regression neural networks
- SOLD: Sinhala Offensive Language Dataset
- DASS: Differentiable Architecture Search for Sparse neural networks
- KDk: A Defense Mechanism Against Label Inference Attacks in Vertical Federated Learning
- Maximizing Discrimination Capability of Knowledge Distillation with Energy Function
- Foundation Model-Based Apple Ripeness and Size Estimation for Selective Harvesting
- MECKD: Deep Learning-Based Fall Detection in Multilayer Mobile Edge Computing With Knowledge Distillation
- Robustness-Reinforced Knowledge Distillation with Correlation Distance and Network Pruning
- PowerSkel: A Device-Free Framework Using CSI Signal for Human Skeleton Estimation in Power Station
- Achieve Fairness without Demographics for Dermatological Disease Diagnosis
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- Deep learning in automated ultrasonic NDE -- developments, axioms and opportunities
- LTD: Low Temperature Distillation for Gradient Masking-free Adversarial Training
- A methodology for training homomorphicencryption friendly neural networks
- Distilling a Deep Neural Network into a Takagi-Sugeno-Kang Fuzzy Inference System
- Finding Deviated Behaviors of the Compressed DNN Models for Image Classifications
- X-Distill: Improving Self-Supervised Monocular Depth via Cross-Task Distillation
- Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
- TIE-KD: Teacher-Independent and Explainable Knowledge Distillation for Monocular Depth Estimation
- Soft Prompt Decoding for Multilingual Dense Retrieval
- Deep Serial Number: Computational Watermarking for DNN Intellectual Property Protection
- NeuroSync: Intent-Aware Code-Based Problem Solving via Direct LLM Understanding Modification
- Distilling Self-Knowledge From Contrastive Links to Classify Graph Nodes Without Passing Messages
- Lightning Fast Video Anomaly Detection via Adversarial Knowledge Distillation
- Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions
- Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference
- Exploring the Mutual Influence between Self-Supervised Single-Frame and Multi-Frame Depth Estimation
- Federated Learning on Non-IID Data: A Survey
- Data Distillation for Text Classification
- ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration
- Reconciling Attribute and Structural Anomalies for Improved Graph Anomaly Detection
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Topological Persistence Guided Knowledge Distillation for Wearable Sensor Data
- Leveraging Angular Distributions for Improved Knowledge Distillation
- Adaptive Intelligence: leveraging insights from adaptive behavior in animals to build flexible AI systems
- Teachers Do More Than Teach: Compressing Image-to-Image Models
- Collaborative Teacher-Student Learning via Multiple Knowledge Transfer
- Mosaicking to Distill: Knowledge Distillation from Out-of-Domain Data
- Keyword localisation in untranscribed speech using visually grounded speech models
- MIME: Adapting a Single Neural Network for Multi-task Inference with Memory-efficient Dynamic Pruning
- Towards Incremental Learning in Large Language Models: A Critical Review
- CORE-ReID: Comprehensive Optimization and Refinement through Ensemble fusion in Domain Adaptation for person re-identification
- Knowledge Distillation on Spatial-Temporal Graph Convolutional Network for Traffic Prediction
- ESAI: Efficient Split Artificial Intelligence via Early Exiting Using Neural Architecture Search
- Prototype Memory for Large-scale Face Representation Learning
- Anti-Distillation: Improving reproducibility of deep networks
- Blockchain-assisted Demonstration Cloning for Multi-Agent Deep Reinforcement Learning
- Foundation Models Knowledge Distillation For Battery Capacity Degradation Forecast
- Learning to Teach with Student Feedback
- Cyber-Physical Systems Security: A Comprehensive Review of Anomaly Detection Techniques
- How and When Adversarial Robustness Transfers in Knowledge Distillation?
- Compressing Neural Networks Using Tensor Networks with Exponentially Fewer Variational Parameters
- LegoNet: Memory Footprint Reduction Through Block Weight Clustering
- Students are the Best Teacher: Exit-Ensemble Distillation with Multi-Exits
- Streaming egocentric action anticipation: An evaluation scheme and approach
- Multi-stage Progressive Compression of Conformer Transducer for On-device Speech Recognition
- DKDL-Net: A Lightweight Bearing Fault Detection Model via Decoupled Knowledge Distillation and Low-Rank Adaptation Fine-tuning
- ImageNet-21K Pretraining for the Masses
- Bidirectional Knowledge Reconfiguration for Lightweight Point Cloud Analysis
- Analyzing Effects of Mixed Sample Data Augmentation on Model Interpretability
- Knowledge Distillation as Semiparametric Inference
- Knowledge Distillation in RNN-Attention Models for Early Prediction of Student Performance
- Improving Neural Topic Models with Wasserstein Knowledge Distillation
- Leveraging Knowledge Distillation for Efficient Deep Reinforcement Learning in Resource-Constrained Environments
- Test-Time Adaptation for Nighttime Color-Thermal Semantic Segmentation
- Scaling Up Quantization-Aware Neural Architecture Search for Efficient Deep Learning on the Edge
- Towards Model Agnostic Federated Learning Using Knowledge Distillation
- MV-MR: multi-views and multi-representations for self-supervised learning and knowledge distillation
- ALPINE: An adaptive language-agnostic pruning method for language models for code
- FedBrain-Distill: Communication-Efficient Federated Brain Tumor Classification Using Ensemble Knowledge Distillation on Non-IID Data
- Response-based Distillation for Incremental Object Detection
- Activation Sparsity Opportunities for Compressing General Large Language Models
- On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards
- I2CKD : Intra- and Inter-Class Knowledge Distillation for Semantic Segmentation
- Growing and Serving Large Open-domain Knowledge Graphs
- ShiftKD: Benchmarking Knowledge Distillation under Distribution Shift
- One Teacher is Enough? Pre-trained Language Model Distillation from Multiple Teachers
- FakeSwarm: Improving Fake News Detection with Swarming Characteristics
- Using Neighborhood Context to Improve Information Extraction from Visual Documents Captured on Mobile Phones
- Robustness and Diversity Seeking Data-Free Knowledge Distillation
- Quantized Convolutional Neural Networks Through the Lens of Partial Differential Equations
- Toward a Unified Framework for Debugging Concept-based Models
- Beyond Self-Supervision: A Simple Yet Effective Network Distillation Alternative to Improve Backbones
- Ground Reaction Force Estimation via Time-aware Knowledge Distillation
- Linear Item-Item Model with Neural Knowledge for Session-based Recommendation
- Selective Knowledge Distillation for Neural Machine Translation
- BeSound: Bluetooth-Based Position Estimation Enhancing with Cross-Modality Distillation
- Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
- A Fast Knowledge Distillation Framework for Visual Recognition
- A Systematic Review of ECG Arrhythmia Classification: Adherence to Standards, Fair Evaluation, and Embedded Feasibility
- RankDVQA-mini: Knowledge Distillation-Driven Deep Video Quality Assessment
- Real-Time Cell Sorting with Scalable In Situ FPGA-Accelerated Deep Learning
- Comparative Analysis of Transformer Models in Disaster Tweet Classification for Public Safety
- GradFreeBits: Gradient Free Bit Allocation for Dynamic Low Precision Neural Networks
- Adaptive Modality Balanced Online Knowledge Distillation for Brain-Eye-Computer based Dim Object Detection
- Understanding the Logit Distributions of Adversarially-Trained Deep Neural Networks
- Tiny is not small enough: High-quality, low-resource facial animation models through hybrid knowledge distillation
- FiGKD: Fine-Grained Knowledge Distillation via High-Frequency Detail Transfer
- LADSG: Label-Anonymized Distillation and Similar Gradient Substitution for Label Privacy in Vertical Federated Learning
- Complementary Relation Contrastive Distillation
- GNN's Uncertainty Quantification using Self-Distillation
- SciceVPR: Stable Cross-Image Correlation Enhanced Model for Visual Place Recognition
- Semi-Online Knowledge Distillation
- An Overview on Generative AI at Scale with Edge-Cloud Computing
- EEG aided boosting of single-lead ECG based sleep staging with Deep Knowledge Distillation
- Adaptive Distillation: Aggregating Knowledge from Multiple Paths for Efficient Distillation
- Emotion recognition in talking-face videos using persistent entropy and neural networks
- A Survey on GAN Acceleration Using Memory Compression Technique
- Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
- KLiNQ: Knowledge Distillation-Assisted Lightweight Neural Network for Qubit Readout on FPGA
- Offline RL With Resource Constrained Online Deployment
- Distributed Edge Inference: an Experimental Study on Multiview Detection
- Faster Molecular Dynamics with Neural Network Potentials via Distilled Multiple Time-Stepping and Non-Conservative Forces
- How to Explain Neural Networks: an Approximation Perspective
- Student-Teacher Learning from Clean Inputs to Noisy Inputs
- Knowledge Distillation for mmWave Beam Prediction Using Sub-6 GHz Channels
- Domain-Agnostic Clustering with Self-Distillation
- Evolving Knowledge Distillation for Lightweight Neural Machine Translation
- Marginal Utility Diminishes: Exploring the Minimum Knowledge for BERT Knowledge Distillation
- HASTE: A Framework for Training-Free, Dynamic, and Steerable Compression of Pre-Trained Convolutional Neural Networks
- OBSeg: Accurate and Fast Instance Segmentation Framework Using Segmentation Foundation Models with Oriented Bounding Box Prompts
- Learning Compatible Embeddings
- TinyML NLP Scheme for Semantic Wireless Sentiment Classification with Privacy Preservation
- Role of Mixup in Topological Persistence Based Knowledge Distillation for Wearable Sensor Data
- Reformulation is All You Need: Addressing Malicious Text Features in DNNs
- Knowledge Distillation from BERT Transformer to Speech Transformer for Intent Classification
- Knowledge Distillation for Intelligent Softwarized Networks: Advances and Open Challenges
- Multi-granularity for knowledge distillation
- Pose-Guided Feature Learning with Knowledge Distillation for Occluded Person Re-Identification
- Embedding Subspace Partitioning for Dynamic Multi-Objective Retrieval
- BERMo: What can BERT learn from ELMo?
- Knowledge Distillation Decision Tree for Unravelling Black-box Machine Learning Models
- EXACT: How to Train Your Accuracy
- Moss: Proxy Model-based Full-Weight Aggregation in Federated Learning with Heterogeneous Models
- Creating Simple, Interpretable Anomaly Detectors for New Physics in Jet Substructure
- Single Snapshot Distillation for Phase Coded Mask Design in Phase Retrieval
- Training BatchNorm Only in Neural Architecture Search and Beyond
- Be Your Own Best Competitor! Multi-Branched Adversarial Knowledge Transfer
- Annealing Knowledge Distillation
- TAGLETS: A System for Automatic Semi-Supervised Learning with Auxiliary Data
- Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression
- Tips and Tricks for Webly-Supervised Fine-Grained Recognition: Learning from the WebFG 2020 Challenge
- On Cross-Layer Alignment for Model Fusion of Heterogeneous Neural Networks
- LISTEN: Lightweight Industrial Sound-representable Transformer for Edge Notification
- Enhancing deep learning models for time series classification via knowledge distillation
- Selective Correlation Based Knowledge Distillation for Ground Reaction Force Estimation
- Intra-class Patch Swap for Self-Distillation