1 paper
Wen-Shu Fan, Xin-Chun Li, De-Chuan Zhan
Knowledge Distillation (KD) could transfer the ``dark knowledge" of a well-performed yet large neural network to a weaker but lightweight one. From the view of output logits and so…