Understanding Back-Translation at Scale
arXiv:1808.09381
Abstract
An effective method to improve neural machine translation with monolingual data is to augment the parallel training corpus with back-translations of target language sentences. This work broadens the understanding of back-translation and investigates a number of methods to generate synthetic source sentences. We find that in all but resource poor settings back-translations obtained via sampling or noised beam outputs are most effective. Our analysis shows that sampling or noisy synthetic data gives a much stronger training signal than data generated by beam or greedy search. We also compare how synthetic data compares to genuine bitext and study various domain effects. Finally, we scale to hundreds of millions of monolingual sentences and achieve a new state of the art of 35 BLEU on the WMT'14 English-German test set.
12 pages; EMNLP 2018
References in corpus (8)
- Sequence to Sequence Learning with Neural Networks
- Neural Machine Translation in Linear Time
- Toward Multilingual Neural Machine Translation with Universal Encoder and Decoder
- Depthwise Separable Convolutions for Neural Machine Translation
- Hierarchical Neural Story Generation
- Weighted Transformer Network for Machine Translation
- Explaining and Generalizing Back-Translation through Wake-Sleep
- Bi-Directional Neural Machine Translation with Synthetic Parallel Data
Cited by in corpus (66)
- Language Models are Few-Shot Learners
- Unsupervised Data Augmentation for Consistency Training
- Leveraging Pre-trained Checkpoints for Sequence Generation Tasks
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- An Overview of Deep Semi-Supervised Learning
- Natural Language Processing Advancements By Deep Learning: A Survey
- Data Augmentation using Pre-trained Transformer Models
- Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
- Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond
- Multiscale Vision Transformers
- Multiple-Attribute Text Style Transfer
- Self-training Improves Pre-training for Natural Language Understanding
- Incorporating BERT into Parallel Sequence Decoding with Adapters
- Text Data Augmentation Made Simple By Leveraging NLP Cloud APIs
- Multi-Graph Transformer for Free-Hand Sketch Recognition
- Establishing Baselines for Text Classification in Low-Resource Languages
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- Self-Explaining Structures Improve NLP Models
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text Classification
- Context in Neural Machine Translation: A Review of Models and Evaluations
- Low Resource Neural Machine Translation: A Benchmark for Five African Languages
- Scratch that! An Evolution-based Adversarial Attack against Neural Networks
- Augmenting Neural Machine Translation with Knowledge Graphs
- Explaining Bayesian Neural Networks
- Unsupervised Neural Machine Translation with Generative Language Models Only
- Neural Machine Translation: A Review of Methods, Resources, and Tools
- Large-scale Pretraining for Neural Machine Translation with Tens of Billions of Sentence Pairs
- Improving Grammatical Error Correction with Machine Translation Pairs
- Diverse Pretrained Context Encodings Improve Document Translation
- End-to-End Whisper to Natural Speech Conversion using Modified Transformer Network
- BiQGEMM: Matrix Multiplication with Lookup Table For Binary-Coding-based Quantized DNNs
- AR: Auto-Repair the Synthetic Data for Neural Machine Translation
- Improving Molecular Design by Stochastic Iterative Target Augmentation
- Improving Non-autoregressive Neural Machine Translation with Monolingual Data
- Multitask Learning for Class-Imbalanced Discourse Classification
- Unsupervised Parallel Corpus Mining on Web Data
- Clustering of Social Media Messages for Humanitarian Aid Response during Crisis
- Parallel Data Augmentation for Formality Style Transfer
- Fast and Simple Mixture of Softmaxes with BPE and Hybrid-LightRNN for Language Generation
- A Survey on Low-Resource Neural Machine Translation
- Evaluating Low-Resource Machine Translation between Chinese and Vietnamese with Back-Translation
- Neural Machine Translation: A Review and Survey
- Improving the Efficiency of Grammatical Error Correction with Erroneous Span Detection and Correction
- Structure-Invariant Testing for Machine Translation
- Memory-Efficient Differentiable Transformer Architecture Search
- Can We Achieve More with Less? Exploring Data Augmentation for Toxic Comment Classification
- Context-gloss Augmentation for Improving Word Sense Disambiguation
- The USTC-NELSLIP Systems for Simultaneous Speech Translation Task at IWSLT 2021
- On the Language Coverage Bias for Neural Machine Translation
- NVIDIA NeMo Neural Machine Translation Systems for English-German and English-Russian News and Biomedical Tasks at WMT21
- Enhancing Pre-trained Language Model with Lexical Simplification
- Iterative Batch Back-Translation for Neural Machine Translation: A Conceptual Model
- Unsupervised Neural Dialect Translation with Commonality and Diversity Modeling
- Generating Diverse Translation by Manipulating Multi-Head Attention
- Converting the Point of View of Messages Spoken to Virtual Assistants
- Detecting ESG topics using domain-specific language models and data augmentation approaches
- Predictions For Pre-training Language Models
- Joint Text and Label Generation for Spoken Language Understanding
- Is artificial data useful for biomedical Natural Language Processing algorithms?
- Reciprocal Supervised Learning Improves Neural Machine Translation
- The NiuTrans System for the WMT21 Efficiency Task
- Empirical Analysis of Korean Public AI Hub Parallel Corpora and in-depth Analysis using LIWC
- Microsoft Research Asia's Systems for WMT19
- Neural Machine Translation with Monte-Carlo Tree Search
- Self-Consistent Models and Values
- Normalization of Input-output Shared Embeddings in Text Generation Models