activity
20172022
most citedAdaptive Gradient Methods with Dynamic Bound of Learning Rate

189 citations · 799 across the 46 of their papers we have counts for

collaborators

84 papers

cs.CL2022

Gradient Knowledge Distillation for Pre-trained Language Models

Lean Wang, Lei Li, Xu Sun

Knowledge distillation (KD) is an effective framework to transfer knowledge from a large-scale teacher to a compact yet well-performing student. Previous KD practices for pre-train…

cs.CL2022

From Mimicking to Integrating: Knowledge Integration for Pre-Trained Language Models

Lei Li, Yankai Lin, Xuancheng Ren +4

Investigating better ways to reuse the released pre-trained language models (PLMs) can significantly reduce the computational cost and the potential environmental side-effects. Thi…

cs.CL20221 cited

Hierarchical Inductive Transfer for Continual Dialogue Learning

Shaoxiong Feng, Xuancheng Ren, Kan Li +1

Pre-trained models have achieved excellent performance on the dialogue task. However, for the continual increase of online chit-chat scenarios, directly fine-tuning these models fo…

cs.CL20211 cited

RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP Models

Wenkai Yang, Yankai Lin, Peng Li +2

Backdoor attacks, which maliciously control a well-trained model's outputs of the instances with specific triggers, are recently shown to be serious threats to the safety of reusin…

cs.LG202147 cited

Topology-Imbalance Learning for Semi-Supervised Node Classification

Deli Chen, Yankai Lin, Guangxiang Zhao +4

The class imbalance problem, as an important issue in learning node representations, has drawn increasing attention from the community. Although the imbalance considered by existin…

cs.CL2021

Dynamic Knowledge Distillation for Pre-trained Language Models

Lei Li, Yankai Lin, Shuhuai Ren +3

Knowledge distillation~(KD) has been proved effective for compressing large-scale pre-trained language models. However, existing methods conduct KD statically, e.g., the student mo…