10 papers
Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge
Hamid Dadkhahi, Firas Trabelsi, Parker Riley +2
Thinking Large Language Models (LLMs) used as judges for pairwise preferences remain noisy at the single-sample level, and common aggregation rules (majority vote, soft self-consis…
TranslateGemma Technical Report
Mara Finkelstein, Isaac Caswell, Tobias Domhan +18
We present TranslateGemma, a suite of open machine translation models based on the Gemma 3 foundation models. To enhance the inherent multilingual capabilities of Gemma 3 for the t…
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
Juraj Juraska, Tobias Domhan, Mara Finkelstein +5
In this paper, we present our submissions to the unified WMT25 Translation Evaluation Shared Task. For the Quality Score Prediction subtask, we create a new generation of MetricX w…
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
Parker Riley, Daniel Deutsch, Mara Finkelstein +3
Human evaluation of machine translation is in an arms race with translation model quality: as our models get better, our evaluation methods need to be improved to ensure that quali…
Generating Difficult-to-Translate Texts
Vilém Zouhar, Wenda Xu, Parker Riley +4
Machine translation benchmarks sourced from the real world are quickly obsoleted, due to most examples being easy for state-of-the-art translation models. This limits the benchmark…
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
Behzad Shayegh, Jan-Thorsten Peter, David Vilar +4
We investigate the tradeoff between adequacy and fluency in machine translation. We show the severity of this tradeoff at the evaluation level and analyze where popular metrics fal…