Teacher's pet: understanding and mitigating biases in distillation
arXiv:2106.10494
Abstract
Knowledge distillation is widely used as a means of improving the performance of a relatively simple student model using the predictions from a complex teacher model. Several works have shown that distillation significantly boosts the student's overall performance; however, are these gains uniform across all data subgroups? In this paper, we show that distillation can harm performance on certain subgroups, e.g., classes with few associated samples. We trace this behaviour to errors made by the teacher distribution being transferred to and amplified by the student model. To mitigate this problem, we present techniques which soften the teacher influence for subgroups where it is less reliable. Experiments on several image classification benchmarks show that these modifications of distillation maintain boost in overall accuracy, while additionally ensuring improvement in subgroup performance.
21 pages, 8 figures
References in corpus (7)
- Distilling the Knowledge in a Neural Network
- Equality of Opportunity in Supervised Learning
- The Devil is in the Tails: Fine-grained Classification in the Wild
- Class-Balanced Loss Based on Effective Number of Samples
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective
- Channel Distillation: Channel-Wise Attention for Knowledge Distillation
- Knowledge Distillation as Semiparametric Inference