7 papers
Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights
Maharaj Brahma, N J Karthika, Rajat Verma +4
Tokenization plays a pivotal role in NLP and is fundamental to training language models. However, existing tokenizers are often skewed towards high-resource languages, limiting the…
SamasÄmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation
N J Karthika, Keerthana Suryanarayanan, Jahanvi Purohit +3
We release SamasÄmayik, a novel, meticulously curated, large-scale Hindi-Sanskrit corpus, comprising 92,196 parallel sentences. Unlike most data available in Sanskrit, which focus…
MorphTok: Morphologically Grounded Tokenization for Indian Languages
Maharaj Brahma, N J Karthika, Atul Singh +5
Tokenization is a crucial step in NLP, especially with the rise of large language models (LLMs), impacting downstream performance, computational cost, and efficiency. Existing LLMs…
LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages
Karthika N J, Krishnakant Bhatt, Ganesh Ramakrishnan +1
Translating technical terms into lexically similar, low-resource Indian languages remains a challenge due to limited parallel data and the complexity of linguistic structures. We p…
LexGen: Domain-aware Multilingual Lexicon Generation
Ayush Maheshwari, Atul Kumar Singh, Karthika NJ +3
Lexicon or dictionary generation across domains has the potential for societal impact, as it can potentially enhance information accessibility for a diverse user base while preserv…
Semantically Cohesive Word Grouping in Indian Languages
N J Karthika, Adyasha Patra, Nagasai Saketh Naidu +3
Indian languages are inflectional and agglutinative and typically follow clause-free word order. The structure of sentences across most major Indian languages are similar when thei…