activity
20242026
collaborators

12 papers

cs.CL2026

Searching the Internet for Challenging Benchmarks at Scale

Wenda Xu, Vilém Zouhar, Parker Riley +3

Many static benchmarks are beginning to saturate: as models rapidly improve, they achieve near-perfect scores on fixed test sets, leaving little headroom to expose genuine model we…

cs.CL2026

When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation

Wenda Xu, Sweta Agrawal, Vilém Zouhar +2

As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM-as-a-benchmark) -- where a model generates test inputs (LLM-as-a-testset) and evaluates o…

cs.CL2026

Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics

Yilin Zhang, Wenda Xu, Zhongtao Liu +2

Quality Estimation (QE) metrics are vital in machine translation for reference-free evaluation and increasingly serve as selection criteria in data filtering and candidate rerankin…

cs.CL2026

TranslateGemma Technical Report

Mara Finkelstein, Isaac Caswell, Tobias Domhan +18

We present TranslateGemma, a suite of open machine translation models based on the Gemma 3 foundation models. To enhance the inherent multilingual capabilities of Gemma 3 for the t…

cs.CL2025

MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task

Juraj Juraska, Tobias Domhan, Mara Finkelstein +5

In this paper, we present our submissions to the unified WMT25 Translation Evaluation Shared Task. For the Quality Score Prediction subtask, we create a new generation of MetricX w…

cs.CL2025

MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation

Parker Riley, Daniel Deutsch, Mara Finkelstein +3

Human evaluation of machine translation is in an arms race with translation model quality: as our models get better, our evaluation methods need to be improved to ensure that quali…