1 citations · 3 across the 4 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2023★ 1 cited
GKD: A General Knowledge Distillation Framework for Large-scale Pre-trained Language Model
Shicheng Tan, Weng Lam Tam, Yuanchun Wang +9
Currently, the reduction in the parameter scale of large-scale pre-trained language models (PLMs) through knowledge distillation has greatly facilitated their widespread deployment…
cs.CL2023
Are Intermediate Layers and Labels Really Necessary? A General Language Model Distillation Method
Shicheng Tan, Weng Lam Tam, Yuanchun Wang +4
The large scale of pre-trained language models poses a challenge for their deployment on various devices, with a growing emphasis on methods to compress these models, particularly…
cs.CL2022★ 1 cited
Coherence-Based Distributed Document Representation Learning for Scientific Documents
Shicheng Tan, Shu Zhao, Yanping Zhang
Distributed document representation is one of the basic problems in natural language processing. Currently distributed document representation methods mainly consider the context i…