5 citations · 7 across the 2 of their papers we have counts for
4 papers
LightMBERT: A Simple Yet Effective Method for Multilingual BERT Distillation
Xiaoqi Jiao, Yichun Yin, Lifeng Shang +5
The multilingual pre-trained language models (e.g, mBERT, XLM and XLM-R) have shown impressive performance on cross-lingual natural language understanding tasks. However, these mod…
Improving Task-Agnostic BERT Distillation with Layer Mapping Search
Xiaoqi Jiao, Huating Chang, Yichun Yin +6
Knowledge distillation (KD) which transfers the knowledge from a large teacher model to a small student model, has been widely used to compress the BERT model recently. Besides the…
TinyBERT: Distilling BERT for Natural Language Understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang +5
Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks. However, pre-trained language models are usually c…
Open Named Entity Modeling from Embedding Distribution
Ying Luo, Hai Zhao, Zhuosheng Zhang +1
In this paper, we report our discovery on named entity distribution in a general word embedding space, which helps an open definition on multilingual named entity definition rather…