1 paper · 1 filter
Yuanda Xu, Hejian Sang, Zhengze Zhou +2
Standard LLM distillation treats all training problems equally -- wasting compute on problems the student has already mastered or cannot yet solve. We empirically show that this in…