Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
arXiv:1907.05019
Abstract
We introduce our efforts towards building a universal neural machine translation (NMT) system capable of translating between any language pair. We set a milestone towards this goal by building a single massively multilingual NMT model handling 103 languages trained on over 25 billion examples. Our system demonstrates effective transfer learning ability, significantly improving translation quality of low-resource languages, while keeping high-resource language translation quality on-par with competitive bilingual baselines. We provide in-depth analysis of various aspects of model building that are crucial to achieving quality and practicality in universal NMT. While we prototype a high-quality universal translation system, our extensive empirical analysis exposes issues that need to be further addressed, and we suggest directions for future research.
References in corpus (19)
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Methods for Interpreting and Understanding Deep Neural Networks
- An Overview of Multi-Task Learning in Deep Neural Networks
- Theoretical Models of Learning to Learn
- On Using Monolingual Corpora in Neural Machine Translation
- Deep Learning Scaling is Predictable, Empirically
- Massively Multitask Networks for Drug Discovery
- Pay Less Attention with Lightweight and Dynamic Convolutions
- Toward Multilingual Neural Machine Translation with Universal Encoder and Decoder
- One Model To Learn Them All
- Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
- Parameter-Efficient Transfer Learning for NLP
- Fixup Initialization: Residual Learning Without Normalization
- The Missing Ingredient in Zero-Shot Neural Machine Translation
- Continual Learning via Neural Pruning
- Learn to Grow: A Continual Structure Learning Framework for Overcoming Catastrophic Forgetting
- SYSTRAN's Pure Neural Machine Translation Systems
- Improved Zero-shot Neural Machine Translation via Ignoring Spurious Correlations
Cited by in corpus (49)
- Multilingual Denoising Pre-training for Neural Machine Translation
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
- Multilingual Translation with Extensible Multilingual Pretraining and Finetuning
- IndicBART: A Pre-trained Model for Indic Natural Language Generation
- Finetuned Language Models Are Zero-Shot Learners
- SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
- Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation
- DAPPLE: A Pipelined Data Parallel Approach for Training Large Models
- A Comprehensive Survey of Multilingual Neural Machine Translation
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task Semantics
- Scaling Laws for Neural Machine Translation
- SQuId: Measuring Speech Naturalness in Many Languages
- Scaling End-to-End Models for Large-Scale Multilingual ASR
- Adaptive Sparse Transformer for Multilingual Translation
- Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
- Duplex Sequence-to-Sequence Learning for Reversible Machine Translation
- Marathi To English Neural Machine Translation With Near Perfect Corpus And Transformers
- Training Multilingual Machine Translation by Alternately Freezing Language-Specific Encoders-Decoders
- The Effect of Domain and Diacritics in Yorùbá-English Neural Machine Translation
- Robust Optimization for Multilingual Translation with Imbalanced Data
- Can Monolingual Pretrained Models Help Cross-Lingual Classification?
- Including Signed Languages in Natural Language Processing
- Lite Training Strategies for Portuguese-English and English-Portuguese Translation
- An Open Dataset and Model for Language Identification
- Neural machine translation, corpus and frugality
- Multi-task Learning for Multilingual Neural Machine Translation
- A Survey on Low-Resource Neural Machine Translation
- Continual Learning in Multilingual NMT via Language-Specific Embeddings
- BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine Translation
- Bilingual Language Modeling, A transfer learning technique for Roman Urdu
- Cross-lingual Inductive Transfer to Detect Offensive Language
- Multilingual Medical Question Answering and Information Retrieval for Rural Health Intelligence Access
- Sicilian Translator: A Recipe for Low-Resource NMT
- Sequential Reptile: Inter-Task Gradient Alignment for Multilingual Learning
- MergeDistill: Merging Pre-trained Language Models using Distillation
- ESPnet-ST IWSLT 2021 Offline Speech Translation System
- Improving Multilingual Translation by Representation and Gradient Regularization
- One Network Fits All? Modular versus Monolithic Task Formulations in Neural Networks
- AfroMT: Pretraining Strategies and Reproducible Benchmarks for Translation of 8 African Languages
- Bandits Don't Follow Rules: Balancing Multi-Facet Machine Translation with Multi-Armed Bandits
- Multilingual Translation via Grafting Pre-trained Language Models
- Multilingual AMR-to-Text Generation
- More Parameters? No Thanks!
- Let Your Heart Speak in its Mother Tongue: Multilingual Captioning of Cardiac Signals
- Distributionally Robust Multilingual Machine Translation
- Competence-based Curriculum Learning for Multilingual Machine Translation
- Self-Learning for Zero Shot Neural Machine Translation
- XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages