most citedComparative Study of Pre-Trained BERT Models for Code-Mixed Hindi-English Data

21 citations · 50 across the 15 of their papers we have counts for

collaborators

15 papers

cs.CL20241 cited

Leveraging Parameter Efficient Training Methods for Low Resource Text Classification: A Case Study in Marathi

Pranita Deshmukh, Nikita Kulkarni, Sanhita Kulkarni +2

With the surge in digital content in low-resource languages, there is an escalating demand for advanced Natural Language Processing (NLP) techniques tailored to these languages. BE…

cs.CL20243 cited

Curating Stopwords in Marathi: A TF-IDF Approach for Improved Text Analysis and Information Retrieval

Rohan Chavan, Gaurav Patil, Vishal Madle +1

Stopwords are commonly used words in a language that are often considered to be of little value in determining the meaning or significance of a document. These words occur frequent…

cs.CL2024

TextGram: Towards a better domain-adaptive pretraining

Sharayu Hiwarkhedkar, Saloni Mittal, Vidula Magdum +4

For green AI, it is crucial to measure and reduce the carbon footprint emitted during the training of large language models. In NLP, performing pre-training on Transformer models r…

cs.CL20242 cited

L3Cube-MahaNews: News-based Short Text and Long Document Classification Datasets in Marathi

Saloni Mittal, Vidula Magdum, Omkar Dhekane +2

The availability of text or topic classification datasets in the low-resource Marathi language is limited, typically consisting of fewer than 4 target labels, with some achieving n…

cs.CL20241 cited

MahaSQuAD: Bridging Linguistic Divides in Marathi Question-Answering

Ruturaj Ghatage, Aditya Kulkarni, Rajlaxmi Patil +2

Question-answering systems have revolutionized information retrieval, but linguistic and cultural boundaries limit their widespread accessibility. This research endeavors to bridge…

cs.CL20241 cited

L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages

Aishwarya Mirashi, Srushti Sonavane, Purva Lingayat +2

In this work, we introduce L3Cube-IndicNews, a multilingual text classification corpus aimed at curating a high-quality dataset for Indian regional languages, with a specific focus…