110 citations · 115 across the 3 of their papers we have counts for
4 papers · 1 filter
Preliminary WMT24 Ranking of General MT Systems and LLMs
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden +18
This is the preliminary ranking of WMT24 General MT systems based on automatic metrics. The official ranking will be a human evaluation, which is superior to the automatic ranking…
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
Nikita Moghe, Arnisa Fazla, Chantal Amrhein +5
Recent machine translation (MT) metrics calibrate their effectiveness by correlating with human judgement but without any insights about their behaviour across different error type…
GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4
Tom Kocmi, Christian Federmann
This paper introduces GEMBA-MQM, a GPT-based evaluation metric designed to detect translation quality errors, specifically for the quality estimation setting without the need for h…
Large Language Models Are State-of-the-Art Evaluators of Translation Quality
Tom Kocmi, Christian Federmann
We describe GEMBA, a GPT-based metric for assessment of translation quality, which works both with a reference translation and without. In our evaluation, we focus on zero-shot pro…