Adaptation of Deep Bidirectional Multilingual Transformers for Russian Language
arXiv:1905.07213
Abstract
The paper introduces methods of adaptation of multilingual masked language models for a specific language. Pre-trained bidirectional language models show state-of-the-art performance on a wide range of tasks including reading comprehension, natural language inference, and sentiment analysis. At the moment there are two alternative approaches to train such models: monolingual and multilingual. While language specific models show superior performance, multilingual models allow to perform a transfer from one language to another and solve tasks for different languages simultaneously. This work shows that transfer learning from a multilingual model to monolingual model results in significant growth of performance on such tasks as reading comprehension, paraphrase detection, and sentiment analysis. Furthermore, multilingual initialization of monolingual model substantially reduces training time. Pre-trained models for the Russian language are open sourced.
Cited by in corpus (22)
- Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- A Systematic Analysis of Morphological Content in BERT Models for Multiple Languages
- Latin BERT: A Contextual Language Model for Classical Philology
- WikiBERT models: deep transfer learning for many languages
- X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language Models
- The ADAPT Enhanced Dependency Parser at the IWPT 2020 Shared Task
- On the Complementarity of Data Selection and Fine Tuning for Domain Adaptation
- Evaluation of contextual embeddings on less-resourced languages
- Detecting Inappropriate Messages on Sensitive Topics that Could Harm a Company's Reputation
- NEREL: A Russian Dataset with Nested Named Entities, Relations and Events
- Transfer Learning for Multi-lingual Tasks -- a Survey
- Bilingual Language Modeling, A transfer learning technique for Roman Urdu
- RUSSE'2020: Findings of the First Taxonomy Enrichment Task for the Russian language
- Neural Coreference Resolution for Arabic
- Text-based classification of interviews for mental health -- juxtaposing the state of the art
- Russian Natural Language Generation: Creation of a Language Modelling Dataset and Evaluation with Modern Neural Architectures
- MergeDistill: Merging Pre-trained Language Models using Distillation
- LOME: Large Ontology Multilingual Extraction
- belabBERT: a Dutch RoBERTa-based language model applied to psychiatric classification
- Evaluating the Efficacy of Summarization Evaluation across Languages
- Domain-Transferable Method for Named Entity Recognition Task