1 paper
Zihe Liu, Yulong Mao, Jinan Xu +2
Knowledge distillation is an effective technique for pre-trained language model compression. However, existing methods only focus on the knowledge distribution among layers, which…