works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CL2026

Dynamically Allocating Evaluation Effort for Model Ranking

Vilém Zouhar, Vilém Zouhar, Julia Kreutzer +6

While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation pr…

cs.CL2026

Contrastive ESA: Human Evaluation of Multiple Translations at Once

Vilém Zouhar, Roman Grundkiewicz, Sara Rajaee +6

The paper proposes Contrastive Error Span Annotation (cESA), a human evaluation protocol that shows multiple translations of the same source together, lets annotators mark error sp…

cs.CL2026

LKValues: Aligning Large Language Models with Sri Lankan Societal Values

Nethmi Muthugala, Supryadi, Surangika Ranathunga +7

Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societi…

cs.CL2026

Pearmut: Human Evaluation of Translation Made Trivial

Vilém Zouhar, Tom Kocmi

Human evaluation is the gold standard for multilingual NLP, but is often skipped in practice and substituted with automatic metrics because it is notoriously complex and slow to se…

cs.CL2026

Unlocking Reasoning Capability on Machine Translation in Large Language Models

Sara Rajaee, Sebastian Vincent, Alexandre Berard +3

Reasoning-oriented large language models (RLMs) achieve strong gains on tasks such as mathematics and coding by generating explicit intermediate reasoning. However, their impact on…

cs.CL2025

Estimating Machine Translation Difficulty

Lorenzo Proietti, Stefano Perrella, Vilém Zouhar +2

Machine translation quality has steadily improved over the years, achieving near-perfect translations in recent benchmarks. These high-quality outputs make it difficult to distingu…