1 paper
Sibgat Ul Islam, Jawad Ibn Ahad, Fuad Rahman +3
Knowledge Distillation (KD) trains a smaller student model using a large, pre-trained teacher model, with temperature as a key hyperparameter controlling the softness of output pro…