1 paper · 1 filter
Zilong Liu, Xuewen Zhang, Jinrui Xing +3
Knowledge distillation (KD) is a key technique for compressing Large Language Models (LLMs), yet methods relying on a single KL objective often fail to balance primary distribution…