activity
20242026
collaborators

8 papers

cs.CL2026

Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights

Maharaj Brahma, N J Karthika, Rajat Verma +4

Tokenization plays a pivotal role in NLP and is fundamental to training language models. However, existing tokenizers are often skewed towards high-resource languages, limiting the…

cs.CL2025

DIWALI: Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context

Pramit Sahoo, Maharaj Brahma, Maunendra Sankar Desarkar

Large language models (LLMs) are widely used in various tasks and applications. However, despite their wide capabilities, they are shown to lack cultural alignment \citep{ryan-etal…

cs.CL2025

MorphTok: Morphologically Grounded Tokenization for Indian Languages

Maharaj Brahma, N J Karthika, Atul Singh +5

Tokenization is a crucial step in NLP, especially with the rise of large language models (LLMs), impacting downstream performance, computational cost, and efficiency. Existing LLMs…

cs.CL2025

AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda

Mohd Nauman, Sravan Gvm, Vijay Devane +7

Current large language models excel at broad, general-purpose tasks, but consistently underperform when exposed to highly specialized domains that require deep cultural, linguistic…

cs.CL2025

The Art of Breaking Words: Rethinking Multilingual Tokenizer Design

Aamod Thakur, Ajay Nagpal, Atharva Savarkar +7

While model architecture and training objectives are well-studied, tokenization, particularly in multilingual contexts, remains a relatively neglected aspect of Large Language Mode…

cs.CL2025

BoK: Introducing Bag-of-Keywords Loss for Interpretable Dialogue Response Generation

Suvodip Dey, Maunendra Sankar Desarkar

The standard language modeling (LM) loss by itself has been shown to be inadequate for effective dialogue modeling. As a result, various training approaches, such as auxiliary loss…