From the 1 of 6 linked papers with an AI index.
6 papers
Contrastive ESA: Human Evaluation of Multiple Translations at Once
Vilém Zouhar, Roman Grundkiewicz, Sara Rajaee +6
The paper proposes Contrastive Error Span Annotation (cESA), a human evaluation protocol that shows multiple translations of the same source together, lets annotators mark error sp…
AI translation of literary texts is "fine", but readers still prefer human translations
Yves Ferstler, Adam Podoxin, Ty Brassington +3
AI translation of literary works is increasingly common. While the content may be rendered adequately, we do not know enough about how readers experience it in terms of immersivene…
PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
Lorenzo Proietti, Roman Grundkiewicz, Matt Post
We present PEAR (Pairwise Evaluation for Automatic Relative Scoring), a supervised quality estimation (QE) metric family that reframes reference-free machine translation (MT) evalu…
Preliminary Ranking of WMT25 General Machine Translation Systems
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden +25
We present the preliminary rankings of machine translation (MT) systems submitted to the WMT25 General Machine Translation Shared Task, as determined by automatic evaluation metric…
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
Tom Kocmi, Vilém Zouhar, Eleftherios Avramidis +5
High-quality Machine Translation (MT) evaluation relies heavily on human judgments. Comprehensive error classification methods, such as Multidimensional Quality Metrics (MQM), are…
On Instruction-Finetuning Neural Machine Translation Models
Vikas Raunak, Roman Grundkiewicz, Marcin Junczys-Dowmunt
In this work, we introduce instruction finetuning for Neural Machine Translation (NMT) models, which distills instruction following capabilities from Large Language Models (LLMs) i…