collaborators

5 papers

cs.CL2026

Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results

Jan-Thorsten Peter, David Vilar, Tobias Domhan +2

Most current large language models (LLMs) support a wide variety of languages in addition to English, including high-resource languages (e.g. German, Chinese, French), as well as l…

cs.CL2026

TranslateGemma Technical Report

Mara Finkelstein, Isaac Caswell, Tobias Domhan +18

We present TranslateGemma, a suite of open machine translation models based on the Gemma 3 foundation models. To enhance the inherent multilingual capabilities of Gemma 3 for the t…

cs.CL2025

MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task

Juraj Juraska, Tobias Domhan, Mara Finkelstein +5

In this paper, we present our submissions to the unified WMT25 Translation Evaluation Shared Task. For the Quality Score Prediction subtask, we create a new generation of MetricX w…

cs.CL2025

Same evaluation, more tokens: On the effect of input length for machine translation evaluation using Large Language Models

Tobias Domhan, Dawei Zhu

Accurately evaluating machine-translated text remains a long-standing challenge, particularly for long documents. Recent work has shown that large language models (LLMs) can serve…

cs.CL2025

Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation

Behzad Shayegh, Jan-Thorsten Peter, David Vilar +4

We investigate the tradeoff between adequacy and fluency in machine translation. We show the severity of this tradeoff at the evaluation level and analyze where popular metrics fal…