1 paper
Yuxian Gu, Li Dong, Furu Wei +1
Knowledge Distillation (KD) is a promising technique for reducing the high computational demand of large language models (LLMs). However, previous KD methods are primarily applied…