2 citations · 2 across the 1 of their papers we have counts for
1 paper
Xiaoqi Jiao, Huating Chang, Yichun Yin +6
Knowledge distillation (KD) which transfers the knowledge from a large teacher model to a small student model, has been widely used to compress the BERT model recently. Besides the…