189 citations · 799 across the 46 of their papers we have counts for
66 papers · 1 filter
Gradient Knowledge Distillation for Pre-trained Language Models
Lean Wang, Lei Li, Xu Sun
Knowledge distillation (KD) is an effective framework to transfer knowledge from a large-scale teacher to a compact yet well-performing student. Previous KD practices for pre-train…
From Mimicking to Integrating: Knowledge Integration for Pre-Trained Language Models
Lei Li, Yankai Lin, Xuancheng Ren +4
Investigating better ways to reuse the released pre-trained language models (PLMs) can significantly reduce the computational cost and the potential environmental side-effects. Thi…
Hierarchical Inductive Transfer for Continual Dialogue Learning
Shaoxiong Feng, Xuancheng Ren, Kan Li +1
Pre-trained models have achieved excellent performance on the dialogue task. However, for the continual increase of online chit-chat scenarios, directly fine-tuning these models fo…
RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP Models
Wenkai Yang, Yankai Lin, Peng Li +2
Backdoor attacks, which maliciously control a well-trained model's outputs of the instances with specific triggers, are recently shown to be serious threats to the safety of reusin…
Dynamic Knowledge Distillation for Pre-trained Language Models
Lei Li, Yankai Lin, Shuhuai Ren +3
Knowledge distillation~(KD) has been proved effective for compressing large-scale pre-trained language models. However, existing methods conduct KD statically, e.g., the student mo…
Text AutoAugment: Learning Compositional Augmentation Policy for Text Classification
Shuhuai Ren, Jinchao Zhang, Lei Li +2
Data augmentation aims to enrich training samples for alleviating the overfitting issue in low-resource or class-imbalanced situations. Traditional methods first devise task-specif…