activity
20242026
collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension

Raviraj Joshi, Utkarsh Vaidya, Sanjay Singh Chauhan +1

Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token embeddings can strongly affe…

cs.CL2025

Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis

Anusha Kamath, Kanishk Singla, Rakesh Paul +4

Evaluating instruction-tuned Large Language Models (LLMs) in Hindi is challenging due to a lack of high-quality benchmarks, as direct translation of English datasets fails to captu…

cs.CL2025

CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications

Raviraj Joshi, Rakesh Paul, Kanishk Singla +8

The increasing use of Large Language Models (LLMs) in agentic applications highlights the need for robust safety guard models. While content safety in English is well-studied, non-…

cs.CL2025

Aligning Large Language Models to Low-Resource Languages through LLM-Based Selective Translation: A Systematic Study

Rakesh Paul, Anusha Kamath, Kanishk Singla +4

Multilingual large language models (LLMs) often demonstrate a performance gap between English and non-English languages, particularly in low-resource settings. Aligning these model…

cs.CL2024

On Importance of Code-Mixed Embeddings for Hate Speech Identification

Shruti Jagdale, Omkar Khade, Gauri Takalikar +2

Code-mixing is the practice of using two or more languages in a single sentence, which often occurs in multilingual communities such as India where people commonly speak multiple l…

cs.CL2024

Challenges in Adapting Multilingual LLMs to Low-Resource Languages using LoRA PEFT Tuning

Omkar Khade, Shruti Jagdale, Abhishek Phaltankar +2

Large Language Models (LLMs) have demonstrated remarkable multilingual capabilities, yet challenges persist in adapting these models for low-resource languages. In this study, we i…