3 papers
cs.CL2025
L3Cube-MahaSTS: A Marathi Sentence Similarity Dataset and Models
Aishwarya Mirashi, Ananya Joshi, Raviraj Joshi
We present MahaSTS, a human-annotated Sentence Textual Similarity (STS) dataset for Marathi, along with MahaSBERT-STS-v2, a fine-tuned Sentence-BERT model optimized for regression-…
cs.CL2024
On Importance of Pruning and Distillation for Efficient Low Resource NLP
Aishwarya Mirashi, Purva Lingayat, Srushti Sonavane +3
The rise of large transformer models has revolutionized Natural Language Processing, leading to significant advances in tasks like text classification. However, this progress deman…
cs.CL2024
L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages
Aishwarya Mirashi, Srushti Sonavane, Purva Lingayat +2
In this work, we introduce L3Cube-IndicNews, a multilingual text classification corpus aimed at curating a high-quality dataset for Indian regional languages, with a specific focus…