AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages
arXiv:2005.00085
Abstract
We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article category classification datasets for 9 languages to evaluate the embeddings. We show that the IndicNLP embeddings significantly outperform publicly available pre-trained embedding on multiple evaluation tasks. We hope that the availability of the corpus will accelerate Indic NLP research. The resources are available at https://github.com/ai4bharat-indicnlp/indicnlp_corpus.
7 pages, 8 tables, https://github.com/ai4bharat-indicnlp/indicnlp_corpus
References in corpus (1)
Cited by in corpus (5)
- The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
- Indic-Transformers: An Analysis of Transformer Language Models for Indian Languages
- Bangla Text Classification using Transformers
- Crosslingual Embeddings are Essential in UNMT for Distant Languages: An English to IndoAryan Case Study
- HinFlair: pre-trained contextual string embeddings for pos tagging and text classification in the Hindi language