1 paper
Animesh Koratana, Daniel Kang, Peter Bailis +1
Knowledge distillation (KD) is a popular method for reducing the computational overhead of deep network inference, in which the output of a teacher model is used to train a smaller…