8 citations · 13 across the 3 of their papers we have counts for
3 papers
cs.CL2021★ 5 cited
Diformer: Directional Transformer for Neural Machine Translation
Minghan Wang, Jiaxin Guo, Yuxia Wang +8
Autoregressive (AR) and Non-autoregressive (NAR) models have their own superiority on the performance and latency, combining them into one model may take advantage of both. Current…
cs.CL2021
Joint-training on Symbiosis Networks for Deep Nueral Machine Translation models
Zhengzhe Yu, Jiaxin Guo, Minghan Wang +11
Deep encoders have been proven to be effective in improving neural machine translation (NMT) systems, but it reaches the upper bound of translation quality when the number of encod…
cs.CL2021★ 8 cited
Self-Distillation Mixup Training for Non-autoregressive Neural Machine Translation
Jiaxin Guo, Minghan Wang, Daimeng Wei +11
Recently, non-autoregressive (NAT) models predict outputs in parallel, achieving substantial improvements in generation speed compared to autoregressive (AT) models. While performi…