1 paper
Guanghui Wang, Kaiwen Lv Kacuila, Zhiyong Yang +5
Knowledge distillation (KD) transfers knowledge from a large teacher model to a smaller student. In language modeling, the student is trained either on tokens sampled from the teac…