Cross-lingual Language Model Pretraining
arXiv:1901.07291
Abstract
Recent studies have demonstrated the efficiency of generative pretraining for English natural language understanding. In this work, we extend this approach to multiple languages and show the effectiveness of cross-lingual pretraining. We propose two methods to learn cross-lingual language models (XLMs): one unsupervised that only relies on monolingual data, and one supervised that leverages parallel data with a new cross-lingual language model objective. We obtain state-of-the-art results on cross-lingual classification, unsupervised and supervised machine translation. On XNLI, our approach pushes the state of the art by an absolute gain of 4.9% accuracy. On unsupervised machine translation, we obtain 34.3 BLEU on WMT'16 German-English, improving the previous state of the art by more than 9 BLEU. On supervised machine translation, we obtain a new state of the art of 38.5 BLEU on WMT'16 Romanian-English, outperforming the previous best approach by more than 4 BLEU. Our code and pretrained models will be made publicly available.
References in corpus (1)
Cited by in corpus (50)
- ERNIE: Enhanced Representation through Knowledge Integration
- Multilingual Denoising Pre-training for Neural Machine Translation
- Evaluating Word Embedding Models: Methods and Experimental Results
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Reducing Transformer Depth on Demand with Structured Dropout
- Cross-Lingual Ability of Multilingual BERT: An Empirical Study
- Multilingual Translation with Extensible Multilingual Pretraining and Finetuning
- Multilingual Alignment of Contextual Word Representations
- BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model
- How multilingual is Multilingual BERT?
- Does BERT Make Any Sense? Interpretable Word Sense Disambiguation with Contextualized Embeddings
- Multilingual is not enough: BERT for Finnish
- Pre-training via Paraphrasing
- GREEK-BERT: The Greeks visiting Sesame Street
- A Mutual Information Maximization Perspective of Language Representation Learning
- Unsupervised Translation of Programming Languages
- XLDA: Cross-Lingual Data Augmentation for Natural Language Inference and Question Answering
- Exploiting BERT for End-to-End Aspect-based Sentiment Analysis
- Gmail Smart Compose: Real-Time Assisted Writing
- A Focus on Neural Machine Translation for African Languages
- Tree-structured Attention with Hierarchical Accumulation
- A Comprehensive Survey of Multilingual Neural Machine Translation
- Knowledge Distillation from Internal Representations
- Exploration of Neural Machine Translation in Autoformalization of Mathematics in Mizar
- Joint Source-Target Self Attention with Locality Constraints
- Effective Cross-lingual Transfer of Neural Machine Translation Models without Shared Vocabularies
- Zero-Shot Paraphrase Generation with Multilingual Language Models
- Finding Experts in Transformer Models
- Attention-Informed Mixed-Language Training for Zero-shot Cross-lingual Task-oriented Dialogue Systems
- Improved Zero-shot Neural Machine Translation via Ignoring Spurious Correlations
- Acquiring Knowledge from Pre-trained Model to Neural Machine Translation
- CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual NLP
- Bilingual Dictionary Based Neural Machine Translation without Using Parallel Sentences
- Cross-lingual Pre-training Based Transfer for Zero-shot Neural Machine Translation
- Large-scale Pretraining for Neural Machine Translation with Tens of Billions of Sentence Pairs
- Cross-Lingual Transfer Learning for Question Answering
- Self-Supervised Dialogue Learning
- Transformers to Learn Hierarchical Contexts in Multiparty Dialogue for Span-based Question Answering
- Can Monolingual Pretrained Models Help Cross-Lingual Classification?
- Unsupervised Parallel Corpus Mining on Web Data
- Unsupervised Pre-training for Natural Language Generation: A Literature Review
- A Survey of Syntactic-Semantic Parsing Based on Constituent and Dependency Structures
- El Departamento de Nosotros: How Machine Translated Corpora Affects Language Models in MRC Tasks
- EmotionGIF-Yankee: A Sentiment Classifier with Robust Model Based Ensemble Methods
- Long-Short Term Masking Transformer: A Simple but Effective Baseline for Document-level Neural Machine Translation
- Learning to Copy for Automatic Post-Editing
- Leveraging Adversarial Training in Self-Learning for Cross-Lingual Text Classification
- Examining the rhetorical capacities of neural language models
- The LIG system for the English-Czech Text Translation Task of IWSLT 2019
- A Simple and Effective Approach to Automatic Post-Editing with Transfer Learning