collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

INDIC DIALECT: A Multi Task Benchmark to Evaluate and Translate in Indian Language Dialects

Tarun Sharma, Manikandan Ravikiran, Sourava Kumar Behera +3

Recent NLP advances focus primarily on standardized languages, leaving most low-resource dialects under-served especially in Indian scenarios. In India, the issue is particularly i…

cs.CL2025

Lexical and Statistical Analysis of Bangla Newspaper and Literature: A Corpus-Driven Study on Diversity, Readability, and NLP Adaptation

Pramit Bhattacharyya, Arnab Bhattacharya

In this paper, we present a comprehensive corpus-driven analysis of Bangla literary and newspaper texts to investigate their lexical diversity, structural complexity and readabilit…

cs.CL2025

A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs

V. S. D. S. Mahesh Akavarapu, Hrishikesh Terdalkar, Pramit Bhattacharyya +5

Large Language Models (LLMs) have demonstrated remarkable generalization capabilities across diverse tasks and languages. In this study, we focus on natural language understanding…

cs.CL2025

BanglaByT5: Byte-Level Modelling for Bangla

Pramit Bhattacharyya, Arnab Bhattacharya

Large language models (LLMs) have achieved remarkable success across various natural language processing tasks. However, most LLM models use traditional tokenizers like BPE and Sen…

cs.CL2025

Semantically Cohesive Word Grouping in Indian Languages

N J Karthika, Adyasha Patra, Nagasai Saketh Naidu +3

Indian languages are inflectional and agglutinative and typically follow clause-free word order. The structure of sentences across most major Indian languages are similar when thei…