3 papers
cs.CL2026
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
Finnur Ãgúst Ingimundarson, Steinunn Rut Friðriksdóttir, Bjarki Ãrmannsson +2
This paper evaluates current Large Language Model (LLM) benchmarking for Icelandic, identifies problems, and calls for improved evaluation methods in low/medium-resource languages…
cs.CL2025
Preliminary Ranking of WMT25 General Machine Translation Systems
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden +25
We present the preliminary rankings of machine translation (MT) systems submitted to the WMT25 General Machine Translation Shared Task, as determined by automatic evaluation metric…
cs.CL2024
Preliminary WMT24 Ranking of General MT Systems and LLMs
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden +18
This is the preliminary ranking of WMT24 General MT systems based on automatic metrics. The official ranking will be a human evaluation, which is superior to the automatic ranking…