PURSUhInT: In Search of Informative Hint Points Based on Layer Clustering for Knowledge Distillation
arXiv:2103.00053 · doi:10.1016/j.eswa.2022.119040
Abstract
One of the most efficient methods for model compression is hint distillation, where the student model is injected with information (hints) from several different layers of the teacher model. Although the selection of hint points can drastically alter the compression performance, conventional distillation approaches overlook this fact and use the same hint points as in the early studies. Therefore, we propose a clustering based hint selection methodology, where the layers of teacher model are clustered with respect to several metrics and the cluster centers are used as the hint points. Our method is applicable for any student network, once it is applied on a chosen teacher network. The proposed approach is validated in CIFAR-100 and ImageNet datasets, using various teacher-student pairs and numerous hint distillation methods. Our results show that hint points selected by our algorithm results in superior compression performance compared to state-of-the-art knowledge distillation algorithms on the same student models and datasets.
Our codes are published on Code Ocean, where the link to our codes is: https://codeocean.com/capsule/4245746/tree/v1
References in corpus (20)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- FitNets: Hints for Thin Deep Nets
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- Similarity of Neural Network Representations Revisited
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability
- Adaptive Multi-Teacher Multi-level Knowledge Distillation
- Contrastive Representation Distillation
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective
- Reducing the Teacher-Student Gap via Spherical Knowledge Disitllation
- Show, Attend and Distill:Knowledge Distillation via Attention-based Feature Matching
- The State of Knowledge Distillation for Classification
- Interactive Knowledge Distillation
- Compressing Deep Neural Networks via Layer Fusion
- Student Network Learning via Evolutionary Knowledge Distillation
- DistPro: Searching A Fast Knowledge Distillation Process via Meta Optimization
- Knowledge Condensation Distillation
- RAIL-KD: RAndom Intermediate Layer Mapping for Knowledge Distillation
- Multi-granularity for knowledge distillation