2 citations · 2 across the 8 of their papers we have counts for
1 paper · 1 filter
Sibgat Ul Islam, Jawad Ibn Ahad, Fuad Rahman +3
Knowledge Distillation (KD) trains a smaller student model using a large, pre-trained teacher model, with temperature as a key hyperparameter controlling the softness of output pro…