9 papers
Source-Free MT Evaluation Is Not MT Evaluation
Baban Gain, Ramakrishna Appicharla, Asif Ekbal
Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments. As a…
Which Tokens Need Context? A Reference-Based Analysis of Translation Responsibility Using Fertility and Entropy
Ramakrishna Appicharla, Baban Gain, Santanu Pal +1
When humans translate, not every word depends equally on the surrounding context. Some tokens, particularly function words like pronouns and auxiliaries, rely heavily on preceding…
Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs
Deepak Kumar, Baban Gain, Asif Ekbal
Automatic Speech Recognition (ASR) transcripts often contain disfluencies, such as fillers, repetitions, and false starts, which reduce readability and hinder downstream applicatio…
One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging
Baban Gain, Asif Ekbal, Trilok Nath Singh
Weight-space model merging combines independently fine-tuned models without accessing original training data, offering a practical alternative to joint training. While merging succ…
Bridging the Linguistic Divide: A Survey on Leveraging Large Language Models for Machine Translation
Baban Gain, Dibyanayan Bandyopadhyay, Asif Ekbal +1
Large Language Models (LLMs) are rapidly reshaping machine translation (MT), particularly by introducing instruction-following, in-context learning, and preference-based alignment…
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
Soham Bhattacharjee, Mukund K Roy, Yathish Poojary +19
India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognize…