collaborators

9 papers

cs.CL2026

Source-Free MT Evaluation Is Not MT Evaluation

Baban Gain, Ramakrishna Appicharla, Asif Ekbal

Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments. As a…

cs.CL2026

Which Tokens Need Context? A Reference-Based Analysis of Translation Responsibility Using Fertility and Entropy

Ramakrishna Appicharla, Baban Gain, Santanu Pal +1

When humans translate, not every word depends equally on the surrounding context. Some tokens, particularly function words like pronouns and auxiliaries, rely heavily on preceding…

cs.CL2026

Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs

Deepak Kumar, Baban Gain, Asif Ekbal

Automatic Speech Recognition (ASR) transcripts often contain disfluencies, such as fillers, repetitions, and false starts, which reduce readability and hinder downstream applicatio…

cs.CL2026

One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging

Baban Gain, Asif Ekbal, Trilok Nath Singh

Weight-space model merging combines independently fine-tuned models without accessing original training data, offering a practical alternative to joint training. While merging succ…

cs.CL2026

Bridging the Linguistic Divide: A Survey on Leveraging Large Language Models for Machine Translation

Baban Gain, Dibyanayan Bandyopadhyay, Asif Ekbal +1

Large Language Models (LLMs) are rapidly reshaping machine translation (MT), particularly by introducing instruction-following, in-context learning, and preference-based alignment…

cs.CL2025

CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems

Soham Bhattacharjee, Mukund K Roy, Yathish Poojary +19

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognize…