activity
20202022
most citedALP-KD: Attention-Based Layer Projection for Knowledge Distillation

8 citations · 8 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CL2022

Triangular Transfer: Freezing the Pivot for Triangular Machine Translation

Meng Zhang, Liangyou Li, Qun Liu

Triangular machine translation is a special case of low-resource machine translation where the language pair of interest has limited parallel data, but both languages have abundant…

cs.CL2021

Uncertainty-Aware Balancing for Multilingual and Multi-Domain Neural Machine Translation Training

Minghao Wu, Yitong Li, Meng Zhang +3

Learning multilingual and multi-domain translation model is challenging as the heterogeneous and imbalanced data make the model converge inconsistently over different corpora in re…

cs.CL20208 cited

ALP-KD: Attention-Based Layer Projection for Knowledge Distillation

Peyman Passban, Yimeng Wu, Mehdi Rezagholizadeh +1

Knowledge distillation is considered as a training and compression strategy in which two neural networks, namely a teacher and a student, are coupled together during training. The…

cs.CL2020

Revisiting Robust Neural Machine Translation: A Transformer Case Study

Peyman Passban, Puneeth S. M. Saladi, Qun Liu

Transformers (Vaswani et al., 2017) have brought a remarkable improvement in the performance of neural machine translation (NMT) systems but they could be surprisingly vulnerable t…

cs.CL2020

Document Graph for Neural Machine Translation

Mingzhou Xu, Liangyou Li, Derek. F. Wong +2

Previous works have shown that contextual information can improve the performance of neural machine translation (NMT). However, most existing document-level NMT methods only consid…