28 citations · 138 across the 14 of their papers we have counts for
1 paper · 1 filter
Mengya Gao, Yujun Shen, Quanquan Li +1
Knowledge distillation (KD) is one of the most potent ways for model compression. The key idea is to transfer the knowledge from a deep teacher model (T) to a shallower student (S)…