1 paper
Xiao Cui, Yulei Qin, Yuting Gao +7
Knowledge distillation (KD) has been widely adopted to compress large language models (LLMs). Existing KD methods investigate various divergence measures including the Kullback-Lei…