activity
20242026
collaborators

5 papers

cs.CL2026

Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights

Maharaj Brahma, N J Karthika, Rajat Verma +4

Tokenization plays a pivotal role in NLP and is fundamental to training language models. However, existing tokenizers are often skewed towards high-resource languages, limiting the…

cs.CL2025

DIWALI: Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context

Pramit Sahoo, Maharaj Brahma, Maunendra Sankar Desarkar

Large language models (LLMs) are widely used in various tasks and applications. However, despite their wide capabilities, they are shown to lack cultural alignment \citep{ryan-etal…

cs.CL2025

MorphTok: Morphologically Grounded Tokenization for Indian Languages

Maharaj Brahma, N J Karthika, Atul Singh +5

Tokenization is a crucial step in NLP, especially with the rise of large language models (LLMs), impacting downstream performance, computational cost, and efficiency. Existing LLMs…

cs.CL2024

NLIP_Lab-IITH Multilingual MT System for WAT24 MT Shared Task

Maharaj Brahma, Pramit Sahoo, Maunendra Sankar Desarkar

This paper describes NLIP Lab's multilingual machine translation system for the WAT24 shared task on multilingual Indic MT task for 22 scheduled languages belonging to 4 language f…

cs.CL2024

NLIP_Lab-IITH Low-Resource MT System for WMT24 Indic MT Shared Task

Pramit Sahoo, Maharaj Brahma, Maunendra Sankar Desarkar

In this paper, we describe our system for the WMT 24 shared task of Low-Resource Indic Language Translation. We consider eng {as, kha, lus, mni} as participating…