Adaptive Sparse Transformer for Multilingual Translation
arXiv:2104.07358
Abstract
Multilingual machine translation has attracted much attention recently due to its support of knowledge transfer among languages and the low cost of training and deployment compared with numerous bilingual models. A known challenge of multilingual models is the negative language interference. In order to enhance the translation quality, deeper and wider architectures are applied to multilingual modeling for larger model capacity, which suffers from the increased inference cost at the same time. It has been pointed out in recent studies that parameters shared among languages are the cause of interference while they may also enable positive transfer. Based on these insights, we propose an adaptive and sparse architecture for multilingual modeling, and train the model to learn shared and language-specific parameters to improve the positive transfer and mitigate the interference. The sparse architecture only activates a sub-network which preserves inference efficiency, and the adaptive design selects different sub-networks based on the input languages. Our model outperforms strong baselines across multiple benchmarks. On the large-scale OPUS dataset with languages, we achieve , and BLEU improvements in one-to-many, many-to-one and zero-shot tasks respectively compared to standard Transformer without increasing the inference cost.
References in corpus (7)
- Multilingual Denoising Pre-training for Neural Machine Translation
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual Models
- Regularizing Deep Multi-Task Networks using Orthogonal Gradients
- Deep Transformers with Latent Depth