1 paper
Maxime Guigon, Lucas Dixon, Michaël E. Sander
Knowledge Distillation (KD) is a critical tool for training Large Language Models (LLMs), yet the majority of research focuses on approaches that rely solely on output logits, negl…