97 citations · 97 across the 1 of their papers we have counts for
1 paper · 1 filter
Jie Song, Ying Chen, Jingwen Ye +1
Knowledge distillation (KD) has become a well established paradigm for compressing deep neural networks. The typical way of conducting knowledge distillation is to train the studen…