Joint Training for Neural Machine Translation Models with Monolingual Data
arXiv:1803.00353
Abstract
Monolingual data have been demonstrated to be helpful in improving translation quality of both statistical machine translation (SMT) systems and neural machine translation (NMT) systems, especially in resource-poor or domain adaptation tasks where parallel data are not rich enough. In this paper, we propose a novel approach to better leveraging monolingual data for neural machine translation by jointly learning source-to-target and target-to-source NMT models for a language pair with a joint EM optimization method. The training process starts with two initial NMT models pre-trained on parallel data for each direction, and these two models are iteratively updated by incrementally decreasing translation losses on training data. In each iteration step, both NMT models are first used to translate monolingual data from one language to the other, forming pseudo-training data of the other NMT model. Then two new NMT models are learnt from parallel data together with the pseudo training data. Both NMT models are expected to be improved and better pseudo-training data can be generated in next step. Experiment results on Chinese-English and English-German translation tasks show that our approach can simultaneously improve translation quality of source-to-target and target-to-source models, significantly outperforming strong baseline systems which are enhanced with monolingual data for model training including back-translation.
Accepted by AAAI 2018
References in corpus (9)
- Neural Machine Translation by Jointly Learning to Align and Translate
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- ADADELTA: An Adaptive Learning Rate Method
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Effective Approaches to Attention-based Neural Machine Translation
- Dual Learning for Machine Translation
- On Using Monolingual Corpora in Neural Machine Translation
- Modeling Coverage for Neural Machine Translation
Cited by in corpus (10)
- Reaching Human-level Performance in Automatic Grammatical Error Correction: An Empirical Study
- Incorporating BERT into Parallel Sequence Decoding with Adapters
- Image Captioning with Very Scarce Supervised Data: Adversarial Semi-Supervised Learning Approach
- Unsupervised Neural Machine Translation with SMT as Posterior Regularization
- Triangular Architecture for Rare Language Translation
- Acquiring Knowledge from Pre-trained Model to Neural Machine Translation
- Bi-Directional Neural Machine Translation with Synthetic Parallel Data
- Improving Neural Machine Translation with Pre-trained Representation
- Language-Independent Representor for Neural Machine Translation
- Iterative Batch Back-Translation for Neural Machine Translation: A Conceptual Model