1 paper · 1 filter
Sibgat Ul Islam, Jawad Ibn Ahad, Fuad Rahman +3
Knowledge Distillation (KD) trains a smaller student model using a large, pre-trained teacher model, with temperature as a key hyperparameter controlling the softness of output pro…