1 paper · 1 filter
Manish Dhakal, Uthman Jinadu, Anjila Budathoki +2
Standard Knowledge Distillation (KD) compresses Large Language Models (LLMs) by optimizing final outputs, yet it typically treats the teacher's intermediate layer's thought process…