5 citations · 5 across the 1 of their papers we have counts for
1 paper
Yuzhu Wang, Lechao Cheng, Manni Duan +3
Knowledge distillation (KD) exploits a large well-trained model (i.e., teacher) to train a small student model on the same dataset for the same task. Treating teacher features as k…