12 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.LG2022★ 12 cited
Asymmetric Temperature Scaling Makes Larger Networks Teach Well Again
Xin-Chun Li, Wen-Shu Fan, Shaoming Song +4
Knowledge Distillation (KD) aims at transferring the knowledge of a well-performed neural network (the {\it teacher}) to a weaker one (the {\it student}). A peculiar phenomenon is…
cs.LG2021★ 1 cited
A New Training Framework for Deep Neural Network
Zhenyan Hou, Wenxuan Fan
Knowledge distillation is the process of transferring the knowledge from a large model to a small model. In this process, the small model learns the generalization ability of the l…