most citedL3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages

1 citations · 2 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL2024

TextGram: Towards a better domain-adaptive pretraining

Sharayu Hiwarkhedkar, Saloni Mittal, Vidula Magdum +4

For green AI, it is crucial to measure and reduce the carbon footprint emitted during the training of large language models. In NLP, performing pre-training on Transformer models r…

cs.CL20241 cited

L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages

Aishwarya Mirashi, Srushti Sonavane, Purva Lingayat +2

In this work, we introduce L3Cube-IndicNews, a multilingual text classification corpus aimed at curating a high-quality dataset for Indian regional languages, with a specific focus…

cs.CL20231 cited

L3Cube-MahaSocialNER: A Social Media based Marathi NER Dataset and BERT models

Harsh Chaudhari, Anuja Patil, Dhanashree Lavekar +2

This work introduces the L3Cube-MahaSocialNER dataset, the first and largest social media dataset specifically designed for Named Entity Recognition (NER) in the Marathi language.…

cs.CL2023

On Significance of Subword tokenization for Low Resource and Efficient Named Entity Recognition: A case study in Marathi

Harsh Chaudhari, Anuja Patil, Dhanashree Lavekar +3

Named Entity Recognition (NER) systems play a vital role in NLP applications such as machine translation, summarization, and question-answering. These systems identify named entiti…

cs.CL2023

mahaNLP: A Marathi Natural Language Processing Library

Vidula Magdum, Omkar Dhekane, Sharayu Hiwarkhedkar +2

We present mahaNLP, an open-source natural language processing (NLP) library specifically built for the Marathi language. It aims to enhance the support for the low-resource Indian…