1 citations · 2 across the 4 of their papers we have counts for
5 papers
TextGram: Towards a better domain-adaptive pretraining
Sharayu Hiwarkhedkar, Saloni Mittal, Vidula Magdum +4
For green AI, it is crucial to measure and reduce the carbon footprint emitted during the training of large language models. In NLP, performing pre-training on Transformer models r…
L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages
Aishwarya Mirashi, Srushti Sonavane, Purva Lingayat +2
In this work, we introduce L3Cube-IndicNews, a multilingual text classification corpus aimed at curating a high-quality dataset for Indian regional languages, with a specific focus…
L3Cube-MahaSocialNER: A Social Media based Marathi NER Dataset and BERT models
Harsh Chaudhari, Anuja Patil, Dhanashree Lavekar +2
This work introduces the L3Cube-MahaSocialNER dataset, the first and largest social media dataset specifically designed for Named Entity Recognition (NER) in the Marathi language.…
On Significance of Subword tokenization for Low Resource and Efficient Named Entity Recognition: A case study in Marathi
Harsh Chaudhari, Anuja Patil, Dhanashree Lavekar +3
Named Entity Recognition (NER) systems play a vital role in NLP applications such as machine translation, summarization, and question-answering. These systems identify named entiti…
mahaNLP: A Marathi Natural Language Processing Library
Vidula Magdum, Omkar Dhekane, Sharayu Hiwarkhedkar +2
We present mahaNLP, an open-source natural language processing (NLP) library specifically built for the Marathi language. It aims to enhance the support for the low-resource Indian…