Duplex Sequence-to-Sequence Learning for Reversible Machine Translation
arXiv:2105.03458
Abstract
Sequence-to-sequence learning naturally has two directions. How to effectively utilize supervision signals from both directions? Existing approaches either require two separate models, or a multitask-learned model but with inferior performance. In this paper, we propose REDER (Reversible Duplex Transformer), a parameter-efficient model and apply it to machine translation. Either end of REDER can simultaneously input and output a distinct language. Thus REDER enables reversible machine translation by simply flipping the input and output ends. Experiments verify that REDER achieves the first success of reversible machine translation, which helps outperform its multitask-trained baselines by up to 1.3 BLEU.
NeurIPS 2021 camera-ready
References in corpus (10)
- NICE: Non-linear Independent Components Estimation
- Dual Learning for Machine Translation
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation
- KERMIT: Generative Insertion-Based Modeling for Sequences
- Semi-Autoregressive Training Improves Mask-Predict Decoding
- Imitation Learning for Non-Autoregressive Neural Machine Translation
- Non-autoregressive Transformer by Position Learning
- Fully Non-autoregressive Neural Machine Translation: Tricks of the Trade