Robust Optimization for Multilingual Translation with Imbalanced Data
arXiv:2104.07639
Abstract
Multilingual models are parameter-efficient and especially effective in improving low-resource languages by leveraging crosslingual transfer. Despite recent advance in massive multilingual translation with ever-growing model and data, how to effectively train multilingual models has not been well understood. In this paper, we show that a common situation in multilingual training, data imbalance among languages, poses optimization tension between high resource and low resource languages where the found multilingual solution is often sub-optimal for low resources. We show that common training method which upsamples low resources can not robustly optimize population loss with risks of either underfitting high resource languages or overfitting low resource ones. Drawing on recent findings on the geometry of loss landscape and its effect on generalization, we propose a principled optimization algorithm, Curvature Aware Task Scaling (CATS), which adaptively rescales gradients from different tasks with a meta objective of guiding multilingual training to low-curvature neighborhoods with uniformly low loss for all languages. We ran experiments on common benchmarks (TED, WMT and OPUS-100) with varying degrees of data imbalance. CATS effectively improved multilingual optimization and as a result demonstrated consistent gains on low resources ( to BLEU) without hurting high resources. In addition, CATS is robust to overparameterization and large batch size training, making it a promising training method for massive multilingual models that truly improve low resource languages.
References in corpus (28)
- Learning Transferable Visual Models From Natural Language Supervision
- An Overview of Multi-Task Learning in Deep Neural Networks
- Cross-lingual Language Model Pretraining
- Scaling Laws for Neural Language Models
- ERNIE: Enhanced Representation through Knowledge Integration
- The Loss Surfaces of Multilayer Networks
- Multilingual Denoising Pre-training for Neural Machine Translation
- mT5: A massively multilingual pre-trained text-to-text transformer
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Qualitatively characterizing neural network optimization problems
- Understanding and Improving Layer Normalization
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- Multilingual Translation with Extensible Multilingual Pretraining and Finetuning
- How multilingual is Multilingual BERT?
- An Empirical Model of Large-Batch Training
- BERT and PALs: Projected Attention Layers for Efficient Adaptation in Multi-Task Learning
- Pretrained Transformers as Universal Computation Engines
- Understanding the Role of Training Regimes in Continual Learning
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual Models
- Massively Multilingual Neural Machine Translation
- The Break-Even Point on Optimization Trajectories of Deep Neural Networks
- Regularizing Deep Multi-Task Networks using Orthogonal Gradients
- Deep Transformers with Latent Depth
- Are All Languages Created Equal in Multilingual BERT?
- Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts Generalization
- Understanding Why Neural Networks Generalize Well Through GSNR of Parameters
- On Negative Interference in Multilingual Models: Findings and A Meta-Learning Treatment