1 paper
Zilong Liu, Xuewen Zhang, Jinrui Xing +3
Knowledge distillation (KD) is a key technique for compressing Large Language Models (LLMs), yet methods relying on a single KL objective often fail to balance primary distribution…