3 papers
cs.CL2026
Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation
Klaudia-Doris Thellmann, Bernhard Stadler, Michael Färber +1
Machine-translated benchmarks are widely used to assess the multilingual capabilities of large language models (LLMs), yet translation errors in these benchmarks remain underexplor…
cs.CL2026
Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite
Klaudia Thellmann, Bernhard Stadler, Michael Färber
Machine-translated benchmark datasets reduce costs and offer scale, but noise, loss of structure, and uneven quality weaken confidence. What matters is not merely whether we can tr…
cs.CL2024
Towards Multilingual LLM Evaluation for European Languages
Klaudia Thellmann, Bernhard Stadler, Michael Fromm +8
The rise of Large Language Models (LLMs) has revolutionized natural language processing across numerous languages and tasks. However, evaluating LLM performance in a consistent and…