works on

From the 1 of 11 linked papers with an AI index.

activity
20242026
collaborators

11 papers

cs.CL2026

Contrastive ESA: Human Evaluation of Multiple Translations at Once

Vilém Zouhar, Roman Grundkiewicz, Sara Rajaee +6

The paper proposes Contrastive Error Span Annotation (cESA), a human evaluation protocol that shows multiple translations of the same source together, lets annotators mark error sp…

cs.LG2026

Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge

Hamid Dadkhahi, Firas Trabelsi, Parker Riley +2

Thinking Large Language Models (LLMs) used as judges for pairwise preferences remain noisy at the single-sample level, and common aggregation rules (majority vote, soft self-consis…

cs.CL2026

Searching the Internet for Challenging Benchmarks at Scale

Wenda Xu, Vilém Zouhar, Parker Riley +3

Many static benchmarks are beginning to saturate: as models rapidly improve, they achieve near-perfect scores on fixed test sets, leaving little headroom to expose genuine model we…

cs.CL2026

TranslateGemma Technical Report

Mara Finkelstein, Isaac Caswell, Tobias Domhan +18

We present TranslateGemma, a suite of open machine translation models based on the Gemma 3 foundation models. To enhance the inherent multilingual capabilities of Gemma 3 for the t…

cs.CL2025

MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation

Parker Riley, Daniel Deutsch, Mara Finkelstein +3

Human evaluation of machine translation is in an arms race with translation model quality: as our models get better, our evaluation methods need to be improved to ensure that quali…

cs.CL2025

Generating Difficult-to-Translate Texts

Vilém Zouhar, Wenda Xu, Parker Riley +4

Machine translation benchmarks sourced from the real world are quickly obsoleted, due to most examples being easy for state-of-the-art translation models. This limits the benchmark…