2 citations · 2 across the 2 of their papers we have counts for
5 papers · 1 filter
GKD: A General Knowledge Distillation Framework for Large-scale Pre-trained Language Model
Shicheng Tan, Weng Lam Tam, Yuanchun Wang +9
Currently, the reduction in the parameter scale of large-scale pre-trained language models (PLMs) through knowledge distillation has greatly facilitated their widespread deployment…
Task-agnostic Distillation of Encoder-Decoder Language Models
Chen Zhang, Yang Yang, Jingang Wang +1
Finetuning pretrained language models (LMs) have enabled appealing performance on a diverse array of tasks. The intriguing task-agnostic property has driven a shifted focus from ta…
Lifting the Curse of Capacity Gap in Distilling Language Models
Chen Zhang, Yang Yang, Jiahao Liu +4
Pretrained language models (LMs) have shown compelling performance on various downstream tasks, but unfortunately they require a tremendous amount of inference compute. Knowledge d…
XPrompt: Exploring the Extreme of Prompt Tuning
Fang Ma, Chen Zhang, Lei Ren +5
Prompt tuning learns soft prompts to condition frozen Pre-trained Language Models (PLMs) for performing downstream tasks in a parameter-efficient manner. While prompt tuning has gr…
CLOWER: A Pre-trained Language Model with Contrastive Learning over Word and Character Representations
Borun Chen, Hongyin Tang, Jiahao Bu +6
Pre-trained Language Models (PLMs) have achieved remarkable performance gains across numerous downstream tasks in natural language understanding. Various Chinese PLMs have been suc…