Self-Learning for Zero Shot Neural Machine Translation
arXiv:2103.05951
Abstract
Neural Machine Translation (NMT) approaches employing monolingual data are showing steady improvements in resource rich conditions. However, evaluations using real-world low-resource languages still result in unsatisfactory performance. This work proposes a novel zero-shot NMT modeling approach that learns without the now-standard assumption of a pivot language sharing parallel data with the zero-shot source and target languages. Our approach is based on three stages: initialization from any pre-trained NMT model observing at least the target language, augmentation of source sides leveraging target monolingual data, and learning to optimize the initial model to the zero-shot pair, where the latter two constitute a self-learning cycle. Empirical findings involving four diverse (in terms of a language family, script and relatedness) zero-shot pairs show the effectiveness of our approach with up to +5.93 BLEU improvement against a supervised bilingual baseline. Compared to unsupervised NMT, consistent improvements are observed even in a domain-mismatch setting, attesting to the usability of our method.
References in corpus (12)
- Sequence to Sequence Learning with Neural Networks
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Cross-lingual Language Model Pretraining
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Toward Multilingual Neural Machine Translation with Universal Encoder and Decoder
- Transfer Learning across Low-Resource, Related Languages for Neural Machine Translation
- The Missing Ingredient in Zero-Shot Neural Machine Translation
- Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation
- Improved Zero-shot Neural Machine Translation via Ignoring Spurious Correlations
- When and Why is Unsupervised Neural Machine Translation Useless?
- Generalized Data Augmentation for Low-Resource Translation
- Cross-lingual Pre-training Based Transfer for Zero-shot Neural Machine Translation