1 paper · 1 filter
Guanghui Wang, Zhiyong Yang, Zitai Wang +3
Knowledge Distillation (KD) transfers knowledge from a large teacher model to a smaller student model by minimizing the divergence between their output distributions, typically usi…