1 paper
Xinyin Ma, Yongliang Shen, Gongfan Fang +3
Large pre-trained transformer-based language models have achieved impressive results on a wide range of NLP tasks. In the past few years, Knowledge Distillation(KD) has become a po…