UzBERT: pretraining a BERT model for Uzbek
arXiv:2108.09814
Abstract
Pretrained language models based on the Transformer architecture have achieved state-of-the-art results in various natural language processing tasks such as part-of-speech tagging, named entity recognition, and question answering. However, no such monolingual model for the Uzbek language is publicly available. In this paper, we introduce UzBERT, a pretrained Uzbek language model based on the BERT architecture. Our model greatly outperforms multilingual BERT on masked language model accuracy. We make the model publicly available under the MIT open-source license.
9 pages, 1 table
References in corpus (7)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Cross-lingual Language Model Pretraining
- How multilingual is Multilingual BERT?
- Multilingual is not enough: BERT for Finnish
- KR-BERT: A Small-Scale Korean-Specific Language Model
- Development of Word Embeddings for Uzbek Language
- Uzbek Cyrillic-Latin-Cyrillic Machine Transliteration