VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and Generation
arXiv:2010.16046
Abstract
Existing work in multilingual pretraining has demonstrated the potential of cross-lingual transferability by training a unified Transformer encoder for multiple languages. However, much of this work only relies on the shared vocabulary and bilingual contexts to encourage the correlation across languages, which is loose and implicit for aligning the contextual representations between languages. In this paper, we plug a cross-attention module into the Transformer encoder to explicitly build the interdependence between languages. It can effectively avoid the degeneration of predicting masked words only conditioned on the context in its own language. More importantly, when fine-tuning on downstream tasks, the cross-attention module can be plugged in or out on-demand, thus naturally benefiting a wider range of cross-lingual tasks, from language understanding to generation. As a result, the proposed cross-lingual model delivers new state-of-the-art results on various cross-lingual understanding tasks of the XTREME benchmark, covering text classification, sequence labeling, question answering, and sentence retrieval. For cross-lingual generation tasks, it also outperforms all existing cross-lingual models and state-of-the-art Transformer variants on WMT14 English-to-German and English-to-French translation datasets, with gains of up to 1~2 BLEU.
Accepted by ACL 2021 (long paper)
References in corpus (15)
- Cross-lingual Language Model Pretraining
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- Unified Language Model Pre-training for Natural Language Understanding and Generation
- On the Cross-lingual Transferability of Monolingual Representations
- Exploring Simple Siamese Representation Learning
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization
- CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
- ReZero is All You Need: Fast Convergence at Large Depth
- XNLI: Evaluating Cross-lingual Sentence Representations
- Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks
- InfoXLM: An Information-Theoretic Framework for Cross-Lingual Language Model Pre-Training
- Very Deep Transformers for Neural Machine Translation
- XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation
- Pre-training Multilingual Neural Machine Translation by Leveraging Alignment Information
- PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned Generation
Cited by in corpus (5)
- mT5: A massively multilingual pre-trained text-to-text transformer
- DeltaLM: Encoder-Decoder Pre-training for Language Generation and Translation by Augmenting Pretrained Multilingual Encoders
- nmT5 -- Is parallel data still relevant for pre-training massively multilingual language models?
- Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word Alignment
- Aligning Cross-lingual Sentence Representations with Dual Momentum Contrast