works on

From the 1 of 15 linked papers with an AI index.

collaborators

15 papers

cs.CL2026

There is No Theoretical Curse of Multilinguality For Embedding Space Structure

Niyati Bafna, Neha Verma, Vilém Zouhar +2

A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment for large-scale language coverage with a multilingual model.…

cs.CL2026

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu +20

Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone…

cs.CL2026

Dynamically Allocating Evaluation Effort for Model Ranking

Vilém Zouhar, Vilém Zouhar, Julia Kreutzer +6

While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation pr…

cs.CL2026

Contrastive ESA: Human Evaluation of Multiple Translations at Once

Vilém Zouhar, Roman Grundkiewicz, Sara Rajaee +6

The paper proposes Contrastive Error Span Annotation (cESA), a human evaluation protocol that shows multiple translations of the same source together, lets annotators mark error sp…

cs.CL2026

Searching the Internet for Challenging Benchmarks at Scale

Wenda Xu, Vilém Zouhar, Parker Riley +3

Many static benchmarks are beginning to saturate: as models rapidly improve, they achieve near-perfect scores on fixed test sets, leaving little headroom to expose genuine model we…

cs.CL2026

When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation

Wenda Xu, Sweta Agrawal, Vilém Zouhar +2

As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM-as-a-benchmark) -- where a model generates test inputs (LLM-as-a-testset) and evaluates o…