activity
20182022
most citedMinimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine Translation

41 citations · 73 across the 10 of their papers we have counts for

collaborators

10 papers

cs.CL2022

One Reference Is Not Enough: Diverse Distillation with Reference Selection for Non-Autoregressive Translation

Chenze Shao, Xuanfu Wu, Yang Feng

Non-autoregressive neural machine translation (NAT) suffers from the multi-modality problem: the source sentence may have multiple correct translations, but the loss function is ca…

cs.CL2022

Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine Translation

Chenze Shao, Yang Feng

Neural networks tend to gradually forget the previously learned knowledge when learning multiple tasks sequentially from dynamic data distributions. This problem is called \textit{…

cs.CL20211 cited

Guiding Teacher Forcing with Seer Forcing for Neural Machine Translation

Yang Feng, Shuhao Gu, Dengji Guo +2

Although teacher forcing has become the main training paradigm for neural machine translation, it usually makes predictions only conditioned on past information, and hence lacks gl…

cs.CL20212 cited

Sequence-Level Training for Non-Autoregressive Neural Machine Translation

Chenze Shao, Yang Feng, Jinchao Zhang +2

In recent years, Neural Machine Translation (NMT) has achieved notable results in various translation tasks. However, the word-by-word generation manner determined by the autoregre…

cs.CL2021

Modeling Coverage for Non-Autoregressive Neural Machine Translation

Yong Shan, Yang Feng, Chenze Shao

Non-Autoregressive Neural Machine Translation (NAT) has achieved significant inference speedup by generating all tokens simultaneously. Despite its high efficiency, NAT usually suf…

cs.CL2020

Generating Diverse Translation from Model Distribution with Dropout

Xuanfu Wu, Yang Feng, Chenze Shao

Despite the improvement of translation quality, neural machine translation (NMT) often suffers from the lack of diversity in its generation. In this paper, we propose to generate d…