1 paper
Guangda Liu, Yiquan Wang, Chengwei Li +6
Attention distillation, which trains one attention distribution to match another by minimizing their Kullback-Leibler (KL) divergence, is widely used in knowledge distillation, mod…