XLM-T: Scaling up Multilingual Machine Translation with Pretrained Cross-lingual Transformer Encoders
arXiv:2012.15547
Abstract
Multilingual machine translation enables a single model to translate between different languages. Most existing multilingual machine translation systems adopt a randomly initialized Transformer backbone. In this work, inspired by the recent success of language model pre-training, we present XLM-T, which initializes the model with an off-the-shelf pretrained cross-lingual Transformer encoder and fine-tunes it with multilingual parallel data. This simple method achieves significant improvements on a WMT dataset with 10 language pairs and the OPUS-100 corpus with 94 pairs. Surprisingly, the method is also effective even upon the strong baseline with back-translation. Moreover, extensive analysis of XLM-T on unsupervised syntactic parsing, word alignment, and multilingual classification explains its effectiveness for machine translation. The code will be at https://aka.ms/xlm-t.
References in corpus (5)
- Multilingual Denoising Pre-training for Neural Machine Translation
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Toward Multilingual Neural Machine Translation with Universal Encoder and Decoder
- Massively Multilingual Neural Machine Translation
Cited by in corpus (5)
- Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey
- DeltaLM: Encoder-Decoder Pre-training for Language Generation and Translation by Augmenting Pretrained Multilingual Encoders
- GTrans: Grouping and Fusing Transformer Layers for Neural Machine Translation
- HLT-MT: High-resource Language-specific Training for Multilingual Neural Machine Translation
- Text Classification for Predicting Multi-level Product Categories