13 papers
Source-Free MT Evaluation Is Not MT Evaluation
Baban Gain, Ramakrishna Appicharla, Asif Ekbal
Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments. As a…
GPUAlert: A Zero-Instrumentation Process-Boundary Monitor for Diagnosing GPU Training-Job Failures
Parv Agarwal, Asif Ekbal
GPU training jobs fail often, roughly two in five on large production clusters, yet the operator typically learns of a failure only by reconnecting hours later. Experiment trackers…
Which Tokens Need Context? A Reference-Based Analysis of Translation Responsibility Using Fertility and Entropy
Ramakrishna Appicharla, Baban Gain, Santanu Pal +1
When humans translate, not every word depends equally on the surrounding context. Some tokens, particularly function words like pronouns and auxiliaries, rely heavily on preceding…
From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability
Dibyanayan Bandyopadhyay, Asif Ekbal
Sparse autoencoders (SAEs) are increasingly used to extract interpretable features from language models (LMs), yet a central question remains: when can an SAE-based explanation be…
One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging
Baban Gain, Asif Ekbal, Trilok Nath Singh
Weight-space model merging combines independently fine-tuned models without accessing original training data, offering a practical alternative to joint training. While merging succ…
Sparse Semantic Dimension as a Generalization Certificate for LLMs
Dibyanayan Bandyopadhyay, Asif Ekbal
Standard statistical learning theory predicts that Large Language Models (LLMs) should overfit because their parameter counts vastly exceed the number of training tokens. Yet, in p…