Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
arXiv:1707.01219
Abstract
Despite deep neural networks have demonstrated extraordinary power in various applications, their superior performances are at expense of high storage and computational costs. Consequently, the acceleration and compression of neural networks have attracted much attention recently. Knowledge Transfer (KT), which aims at training a smaller student network by transferring knowledge from a larger teacher model, is one of the popular solutions. In this paper, we propose a novel knowledge transfer method by treating it as a distribution matching problem. Particularly, we match the distributions of neuron selectivity patterns between teacher and student networks. To achieve this goal, we devise a new KT loss function by minimizing the Maximum Mean Discrepancy (MMD) metric between these distributions. Combined with the original loss function, our method can significantly improve the performance of student networks. We validate the effectiveness of our method across several datasets, and further combine it with other KT methods to explore the best possible results. Last but not least, we fine-tune the model to other tasks such as object detection. The results are also encouraging, which confirm the transferability of the learned features.
References in corpus (7)
- Distilling the Knowledge in a Neural Network
- Deep Domain Confusion: Maximizing for Domain Invariance
- FitNets: Hints for Thin Deep Nets
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Learning Structured Sparsity in Deep Neural Networks
- Adversarial Discriminative Domain Adaptation
- Channel Pruning for Accelerating Very Deep Neural Networks
Cited by in corpus (39)
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks
- SEED: Self-supervised Distillation For Visual Representation
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Relational Knowledge Distillation
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective
- Channel Distillation: Channel-Wise Attention for Knowledge Distillation
- Variational Information Distillation for Knowledge Transfer
- Distillation Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network
- FEED: Feature-level Ensemble for Knowledge Distillation
- DarkRank: Accelerating Deep Metric Learning via Cross Sample Similarities Transfer
- Knowledge Adaptation for Efficient Semantic Segmentation
- Residual Knowledge Distillation
- Full-Cycle Energy Consumption Benchmark for Low-Carbon Computer Vision
- Learning Student Networks via Feature Embedding
- A Convolutional Neural Network-Based Low Complexity Filter
- MarginDistillation: distillation for margin-based softmax
- Weakly Supervised 3D Object Detection from Point Clouds
- Bidirectional Knowledge Reconfiguration for Lightweight Point Cloud Analysis
- Semantically-Conditioned Negative Samples for Efficient Contrastive Learning
- Triplet Distillation for Deep Face Recognition
- Deep Neural Compression Via Concurrent Pruning and Self-Distillation
- Revisiting Knowledge Distillation: An Inheritance and Exploration Framework
- VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer
- Learning from a Lightweight Teacher for Efficient Knowledge Distillation
- Voice2Mesh: Cross-Modal 3D Face Model Generation from Voices
- Distilling Pixel-Wise Feature Similarities for Semantic Segmentation
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge Distillation
- A Fast Knowledge Distillation Framework for Visual Recognition
- Similarity Transfer for Knowledge Distillation
- Adaptive Distillation: Aggregating Knowledge from Multiple Paths for Efficient Distillation
- Knowledge Representing: Efficient, Sparse Representation of Prior Knowledge for Knowledge Distillation
- Complementary Relation Contrastive Distillation
- Self-Supervised Adaptation for Video Super-Resolution
- Partial to Whole Knowledge Distillation: Progressive Distilling Decomposed Knowledge Boosts Student Better
- CompConv: A Compact Convolution Module for Efficient Feature Learning
- AUTOKD: Automatic Knowledge Distillation Into A Student Architecture Family
- Diversified Mutual Learning for Deep Metric Learning
- A Studious Approach to Semi-Supervised Learning
- Multi-granularity for knowledge distillation