5 papers
Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi
Manasi Waghe, Danish Chandargi, Mohammad Aamir Rayyan +2
Government documents in India are predominantly issued in regional languages such as Marathi, creating substantial accessibility barriers for non-native readers, interstate adminis…
L3Cube-MahaEmotions: A Marathi Emotion Recognition Dataset with Synthetic Annotations using CoTR prompting and Large Language Models
Nidhi Kowtal, Raviraj Joshi
Emotion recognition in low-resource languages like Marathi remains challenging due to limited annotated data. We present L3Cube-MahaEmotions, a high-quality Marathi emotion recogni…
Chain-of-Translation Prompting (CoTR): A Novel Prompting Technique for Low Resource Languages
Tejas Deshpande, Nidhi Kowtal, Raviraj Joshi
This paper introduces Chain of Translation Prompting (CoTR), a novel strategy designed to enhance the performance of language models in low-resource languages. CoTR restructures pr…
Long Range Named Entity Recognition for Marathi Documents
Pranita Deshmukh, Nikita Kulkarni, Sanhita Kulkarni +3
The demand for sophisticated natural language processing (NLP) methods, particularly Named Entity Recognition (NER), has increased due to the exponential growth of Marathi-language…
L3Cube-MahaSum: A Comprehensive Dataset and BART Models for Abstractive Text Summarization in Marathi
Pranita Deshmukh, Nikita Kulkarni, Sanhita Kulkarni +2
We present the MahaSUM dataset, a large-scale collection of diverse news articles in Marathi, designed to facilitate the training and evaluation of models for abstractive summariza…