3 papers
cs.CL2025
Comparative Study of Pre-Trained BERT and Large Language Models for Code-Mixed Named Entity Recognition
Mayur Shirke, Amey Shembade, Pavan Thorat +2
Named Entity Recognition (NER) in code-mixed text, particularly Hindi-English (Hinglish), presents unique challenges due to informal structure, transliteration, and frequent langua…
cs.CL2025
L3Cube-MahaSTS: A Marathi Sentence Similarity Dataset and Models
Aishwarya Mirashi, Ananya Joshi, Raviraj Joshi
We present MahaSTS, a human-annotated Sentence Textual Similarity (STS) dataset for Marathi, along with MahaSBERT-STS-v2, a fine-tuned Sentence-BERT model optimized for regression-…
cs.CL2025
On Importance of Layer Pruning for Smaller BERT Models and Low Resource Languages
Mayur Shirke, Amey Shembade, Madhushri Wagh +2
This study explores the effectiveness of layer pruning for developing more efficient BERT models tailored to specific downstream tasks in low-resource languages. Our primary object…