2 citations · 2 across the 1 of their papers we have counts for
1 paper
Siyue Wu, Hongzhan Chen, Xiaojun Quan +2
Knowledge distillation has attracted a great deal of interest recently to compress pre-trained language models. However, existing knowledge distillation methods suffer from two lim…