1 paper
Shiva Kumar C, Jitendra Kumar Dhiman, Nagaraj Adiga +1
Traditionally, Knowledge Distillation (KD) is used for model compression, often leading to suboptimal performance. In this paper, we evaluate the impact of combining KD loss with a…