1 paper
Yihan Cao, Yanbin Kang, Zhengming Xing +1
Knowledge distillation (KD) is a widely adopted approach for compressing large neural networks by transferring knowledge from a large teacher model to a smaller student model. In t…