activity
20182021
most citedModeling Semantic Compositionality with Sememe Knowledge

7 citations · 18 across the 4 of their papers we have counts for

collaborators

8 papers

cs.CL20214 cited

Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation

Cheng Chen, Yichun Yin, Lifeng Shang +4

Task-agnostic knowledge distillation, a teacher-student framework, has been proved effective for BERT compression. Although achieving promising results on NLP tasks, it requires en…

cs.CL20215 cited

LightMBERT: A Simple Yet Effective Method for Multilingual BERT Distillation

Xiaoqi Jiao, Yichun Yin, Lifeng Shang +5

The multilingual pre-trained language models (e.g, mBERT, XLM and XLM-R) have shown impressive performance on cross-lingual natural language understanding tasks. However, these mod…

cs.CL20202 cited

Improving Task-Agnostic BERT Distillation with Layer Mapping Search

Xiaoqi Jiao, Huating Chang, Yichun Yin +6

Knowledge distillation (KD) which transfers the knowledge from a large teacher model to a small student model, has been widely used to compress the BERT model recently. Besides the…

cs.CL2020

TernaryBERT: Distillation-aware Ultra-low Bit BERT

Wei Zhang, Lu Hou, Yichun Yin +4

Transformer-based pre-training models like BERT have achieved remarkable performance in many natural language processing tasks.However, these models are both computation and memory…

cs.CL2019

A General Framework for Adaptation of Neural Machine Translation to Simultaneous Translation

Yun Chen, Liangyou Li, Xin Jiang +2

Despite the success of neural machine translation (NMT), simultaneous neural machine translation (SNMT), the task of translating in real time before a full sentence has been observ…

cs.CL2019

TinyBERT: Distilling BERT for Natural Language Understanding

Xiaoqi Jiao, Yichun Yin, Lifeng Shang +5

Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks. However, pre-trained language models are usually c…