4 papers
BhashaBench V1: A Comprehensive Benchmark for the Quadrant of Indic Domains
Vijay Devane, Mohd Nauman, Bhargav Patel +14
The rapid advancement of large language models(LLMs) has intensified the need for domain and culture specific evaluation. Existing benchmarks are largely Anglocentric and domain-ag…
The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
Aamod Thakur, Ajay Nagpal, Atharva Savarkar +7
While model architecture and training objectives are well-studied, tokenization, particularly in multilingual contexts, remains a relatively neglected aspect of Large Language Mode…
Intent Aware Context Retrieval for Multi-Turn Agricultural Question Answering
Abhay Vijayvargia, Ajay Nagpal, Kundeshwar Pundalik +5
Indian farmers often lack timely, accessible, and language-friendly agricultural advice, especially in rural areas with low literacy. To address this gap in accessibility, this pap…
PARAM-1 BharatGen 2.9B Model
Kundeshwar Pundalik, Piyush Sawarkar, Nihar Sahoo +19
Large Language Models (LLMs) have emerged as powerful general-purpose reasoning systems, yet their development remains dominated by English-centric data, architectures, and optimiz…