5 papers
Searching the Internet for Challenging Benchmarks at Scale
Wenda Xu, Vilém Zouhar, Parker Riley +3
Many static benchmarks are beginning to saturate: as models rapidly improve, they achieve near-perfect scores on fixed test sets, leaving little headroom to expose genuine model we…
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
Wenda Xu, Sweta Agrawal, Vilém Zouhar +2
As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM-as-a-benchmark) -- where a model generates test inputs (LLM-as-a-testset) and evaluates o…
Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics
Yilin Zhang, Wenda Xu, Zhongtao Liu +2
Quality Estimation (QE) metrics are vital in machine translation for reference-free evaluation and increasingly serve as selection criteria in data filtering and candidate rerankin…
TranslateGemma Technical Report
Mara Finkelstein, Isaac Caswell, Tobias Domhan +18
We present TranslateGemma, a suite of open machine translation models based on the Gemma 3 foundation models. To enhance the inherent multilingual capabilities of Gemma 3 for the t…
Generating Difficult-to-Translate Texts
Vilém Zouhar, Wenda Xu, Parker Riley +4
Machine translation benchmarks sourced from the real world are quickly obsoleted, due to most examples being easy for state-of-the-art translation models. This limits the benchmark…