Contrastive Representation Distillation
arXiv:1910.10699
Abstract
Often we wish to transfer representational knowledge from one neural network to another. Examples include distilling a large network into a smaller one, transferring knowledge from one sensory modality to a second, or ensembling a collection of models into a single estimator. Knowledge distillation, the standard approach to these problems, minimizes the KL divergence between the probabilistic outputs of a teacher and student network. We demonstrate that this objective ignores important structural knowledge of the teacher network. This motivates an alternative objective by which we train a student to capture significantly more information in the teacher's representation of the data. We formulate this objective as contrastive learning. Experiments demonstrate that our resulting new objective outperforms knowledge distillation and other cutting-edge distillers on a variety of knowledge transfer tasks, including single model compression, ensemble distillation, and cross-modal transfer. Our method sets a new state-of-the-art in many transfer tasks, and sometimes even outperforms the teacher network when combined with knowledge distillation. Code: http://github.com/HobbitLong/RepDistiller.
ICLR 2020. Project Page: http://hobbitlong.github.io/CRD/, Code: http://github.com/HobbitLong/RepDistiller. Typo fixed in the newest version
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- FitNets: Hints for Thin Deep Nets
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- On distinguishability criteria for estimating generative models
- Correlation Congruence for Knowledge Distillation
Cited by in corpus (30)
- Ensemble Distillation for Robust Model Fusion in Federated Learning
- Rethinking Few-Shot Image Classification: a Good Embedding Is All You Need?
- Pixel-Level Cycle Association: A New Perspective for Domain Adaptive Semantic Segmentation
- ResKD: Residual-Guided Knowledge Distillation
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Hybrid Discriminative-Generative Training via Contrastive Learning
- Multiple Domain Experts Collaborative Learning: Multi-Source Domain Generalization For Person Re-Identification
- Distilling Knowledge via Knowledge Review
- The State of Knowledge Distillation for Classification
- Mosaicking to Distill: Knowledge Distillation from Out-of-Domain Data
- Full-Cycle Energy Consumption Benchmark for Low-Carbon Computer Vision
- PURSUhInT: In Search of Informative Hint Points Based on Layer Clustering for Knowledge Distillation
- Contrastive Learning for Local and Global Learning MRI Reconstruction
- Contrastive Distillation on Intermediate Representations for Language Model Compression
- Knowledge Distillation for Multi-task Learning
- Semantically-Conditioned Negative Samples for Efficient Contrastive Learning
- Simon Says: Evaluating and Mitigating Bias in Pruned Neural Networks with Knowledge Distillation
- Spirit Distillation: A Model Compression Method with Multi-domain Knowledge Transfer
- Representation Transfer by Optimal Transport
- Deep Neural Compression Via Concurrent Pruning and Self-Distillation
- Multi-view Contrastive Learning for Online Knowledge Distillation
- Follow Your Path: a Progressive Method for Knowledge Distillation
- Multi-level Knowledge Distillation via Knowledge Alignment and Correlation
- Complementary Relation Contrastive Distillation
- Categorical Relation-Preserving Contrastive Knowledge Distillation for Medical Image Classification
- A New Training Framework for Deep Neural Network
- Prime-Aware Adaptive Distillation
- Partial to Whole Knowledge Distillation: Progressive Distilling Decomposed Knowledge Boosts Student Better
- C-SL: Contrastive Sound Localization with Inertial-Acoustic Sensors
- Impression Space from Deep Template Network