21 citations · 21 across the 1 of their papers we have counts for
1 paper
Philip de Rijk, Lukas Schneider, Marius Cordts +1
Knowledge Distillation (KD) is a well-known training paradigm in deep neural networks where knowledge acquired by a large teacher model is transferred to a small student. KD has pr…