5 citations · 7 across the 2 of their papers we have counts for
3 papers
cs.CL2021★ 5 cited
LightMBERT: A Simple Yet Effective Method for Multilingual BERT Distillation
Xiaoqi Jiao, Yichun Yin, Lifeng Shang +5
The multilingual pre-trained language models (e.g, mBERT, XLM and XLM-R) have shown impressive performance on cross-lingual natural language understanding tasks. However, these mod…
cs.CL2020★ 2 cited
Improving Task-Agnostic BERT Distillation with Layer Mapping Search
Xiaoqi Jiao, Huating Chang, Yichun Yin +6
Knowledge distillation (KD) which transfers the knowledge from a large teacher model to a small student model, has been widely used to compress the BERT model recently. Besides the…
cs.CL2019
TinyBERT: Distilling BERT for Natural Language Understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang +5
Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks. However, pre-trained language models are usually c…