collaborators

7 papers

cs.CL2026

Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights

Maharaj Brahma, N J Karthika, Rajat Verma +4

Tokenization plays a pivotal role in NLP and is fundamental to training language models. However, existing tokenizers are often skewed towards high-resource languages, limiting the…

cs.CV2026

Improving Video Question Answering through query-based frame selection

Himanshu Patil, Geo Jolly, Ramana Raja Buddala +2

Video Question Answering (VideoQA) models enhance understanding and interaction with audiovisual content, making it more accessible, searchable, and useful for a wide range of fiel…

cs.CL2025

MorphTok: Morphologically Grounded Tokenization for Indian Languages

Maharaj Brahma, N J Karthika, Atul Singh +5

Tokenization is a crucial step in NLP, especially with the rise of large language models (LLMs), impacting downstream performance, computational cost, and efficiency. Existing LLMs…

cs.CL2025

AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda

Mohd Nauman, Sravan Gvm, Vijay Devane +7

Current large language models excel at broad, general-purpose tasks, but consistently underperform when exposed to highly specialized domains that require deep cultural, linguistic…

cs.CL2025

The Art of Breaking Words: Rethinking Multilingual Tokenizer Design

Aamod Thakur, Ajay Nagpal, Atharva Savarkar +7

While model architecture and training objectives are well-studied, tokenization, particularly in multilingual contexts, remains a relatively neglected aspect of Large Language Mode…

cs.CL2025

Intent Aware Context Retrieval for Multi-Turn Agricultural Question Answering

Abhay Vijayvargia, Ajay Nagpal, Kundeshwar Pundalik +5

Indian farmers often lack timely, accessible, and language-friendly agricultural advice, especially in rural areas with low literacy. To address this gap in accessibility, this pap…