1 paper
Shizhen Li, Zhiyu Shen, Yuyin Lu +4
Knowledge distillation (KD) is widely used to transfer the capabilities of large language models (LLMs) to smaller students, but existing objectives often struggle to balance faith…