Investigating Backtranslation in Neural Machine Translation
arXiv:1804.06189
Abstract
A prerequisite for training corpus-based machine translation (MT) systems -- either Statistical MT (SMT) or Neural MT (NMT) -- is the availability of high-quality parallel data. This is arguably more important today than ever before, as NMT has been shown in many studies to outperform SMT, but mostly when large parallel corpora are available; in cases where data is limited, SMT can still outperform NMT. Recently researchers have shown that back-translating monolingual data can be used to create synthetic parallel corpora, which in turn can be used in combination with authentic parallel data to train a high-quality NMT system. Given that large collections of new parallel text become available only quite rarely, backtranslation has become the norm when building state-of-the-art NMT systems, especially in resource-poor scenarios. However, we assert that there are many unknown factors regarding the actual effects of back-translated data on the translation capabilities of an NMT model. Accordingly, in this work we investigate how using back-translated data as a training corpus -- both as a separate standalone dataset as well as combined with human-generated parallel data -- affects the performance of an NMT model. We use incrementally larger amounts of back-translated data to train a range of NMT systems for German-to-English, and analyse the resulting translation performance.
References in corpus (2)
Cited by in corpus (24)
- CheXbert: Combining Automatic Labelers and Expert Annotations for Accurate Radiology Report Labeling Using BERT
- Corpus Augmentation by Sentence Segmentation for Low-Resource Neural Machine Translation
- Sequence-to-sequence Pre-training with Data Augmentation for Sentence Rewriting
- Simplify-then-Translate: Automatic Preprocessing for Black-Box Machine Translation
- Acquiring Knowledge from Pre-trained Model to Neural Machine Translation
- Improving Back-Translation with Uncertainty-based Confidence Estimation
- Neural Machine Translation: A Review of Methods, Resources, and Tools
- On the Integration of LinguisticFeatures into Statistical and Neural Machine Translation
- Tag-less Back-Translation
- Improving Neural Machine Translation with Pre-trained Representation
- Learning Credible Deep Neural Networks with Rationale Regularization
- AR: Auto-Repair the Synthetic Data for Neural Machine Translation
- A Survey on Low-Resource Neural Machine Translation
- Evaluating Low-Resource Machine Translation between Chinese and Vietnamese with Back-Translation
- Neural Machine Translation: A Review and Survey
- Iterative Batch Back-Translation for Neural Machine Translation: A Conceptual Model
- Generating Diverse Translation by Manipulating Multi-Head Attention
- Pronoun-Targeted Fine-tuning for NMT with Hybrid Losses
- Attentive fine-tuning of Transformers for Translation of low-resourced languages @LoResMT 2021
- Self-Learning for Zero Shot Neural Machine Translation
- Do all Roads Lead to Rome? Understanding the Role of Initialization in Iterative Back-Translation
- The ADAPT System Description for the IWSLT 2018 Basque to English Translation Task
- Language Model-Driven Unsupervised Neural Machine Translation
- Is artificial data useful for biomedical Natural Language Processing algorithms?