8 papers
Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights
Maharaj Brahma, N J Karthika, Rajat Verma +4
Tokenization plays a pivotal role in NLP and is fundamental to training language models. However, existing tokenizers are often skewed towards high-resource languages, limiting the…
DIWALI: Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context
Pramit Sahoo, Maharaj Brahma, Maunendra Sankar Desarkar
Large language models (LLMs) are widely used in various tasks and applications. However, despite their wide capabilities, they are shown to lack cultural alignment \citep{ryan-etal…
MorphTok: Morphologically Grounded Tokenization for Indian Languages
Maharaj Brahma, N J Karthika, Atul Singh +5
Tokenization is a crucial step in NLP, especially with the rise of large language models (LLMs), impacting downstream performance, computational cost, and efficiency. Existing LLMs…
AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda
Mohd Nauman, Sravan Gvm, Vijay Devane +7
Current large language models excel at broad, general-purpose tasks, but consistently underperform when exposed to highly specialized domains that require deep cultural, linguistic…
The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
Aamod Thakur, Ajay Nagpal, Atharva Savarkar +7
While model architecture and training objectives are well-studied, tokenization, particularly in multilingual contexts, remains a relatively neglected aspect of Large Language Mode…
BoK: Introducing Bag-of-Keywords Loss for Interpretable Dialogue Response Generation
Suvodip Dey, Maunendra Sankar Desarkar
The standard language modeling (LM) loss by itself has been shown to be inadequate for effective dialogue modeling. As a result, various training approaches, such as auxiliary loss…