6 citations · 10 across the 4 of their papers we have counts for
4 papers
bert2BERT: Towards Reusable Pretrained Language Models
Cheng Chen, Yichun Yin, Lifeng Shang +7
In recent years, researchers tend to pre-train ever-larger language models to explore the upper limit of deep models. However, large language model pre-training costs intensive com…
Improving Punctuation Restoration for Speech Transcripts via External Data
Xue-Yong Fu, Cheng Chen, Md Tahmid Rahman Laskar +2
Automatic Speech Recognition (ASR) systems generally do not produce punctuated transcripts. To make transcripts more readable and follow the expected input format for downstream la…
AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language Models
Yichun Yin, Cheng Chen, Lifeng Shang +3
Pre-trained language models (PLMs) have achieved great success in natural language processing. Most of PLMs follow the default setting of architecture hyper-parameters (e.g., the h…
Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation
Cheng Chen, Yichun Yin, Lifeng Shang +4
Task-agnostic knowledge distillation, a teacher-student framework, has been proved effective for BERT compression. Although achieving promising results on NLP tasks, it requires en…