5 papers · 1 filter
INDIC DIALECT: A Multi Task Benchmark to Evaluate and Translate in Indian Language Dialects
Tarun Sharma, Manikandan Ravikiran, Sourava Kumar Behera +3
Recent NLP advances focus primarily on standardized languages, leaving most low-resource dialects under-served especially in Indian scenarios. In India, the issue is particularly i…
Lexical and Statistical Analysis of Bangla Newspaper and Literature: A Corpus-Driven Study on Diversity, Readability, and NLP Adaptation
Pramit Bhattacharyya, Arnab Bhattacharya
In this paper, we present a comprehensive corpus-driven analysis of Bangla literary and newspaper texts to investigate their lexical diversity, structural complexity and readabilit…
A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
V. S. D. S. Mahesh Akavarapu, Hrishikesh Terdalkar, Pramit Bhattacharyya +5
Large Language Models (LLMs) have demonstrated remarkable generalization capabilities across diverse tasks and languages. In this study, we focus on natural language understanding…
BanglaByT5: Byte-Level Modelling for Bangla
Pramit Bhattacharyya, Arnab Bhattacharya
Large language models (LLMs) have achieved remarkable success across various natural language processing tasks. However, most LLM models use traditional tokenizers like BPE and Sen…
Semantically Cohesive Word Grouping in Indian Languages
N J Karthika, Adyasha Patra, Nagasai Saketh Naidu +3
Indian languages are inflectional and agglutinative and typically follow clause-free word order. The structure of sentences across most major Indian languages are similar when thei…