3 papers
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
Finnur Ãgúst Ingimundarson, Steinunn Rut Friðriksdóttir, Bjarki Ãrmannsson +2
This paper evaluates current Large Language Model (LLM) benchmarking for Icelandic, identifies problems, and calls for improved evaluation methods in low/medium-resource languages…
Killing Two Flies with One Stone: An Attempt to Break LLMs Using English->Icelandic Idioms and Proper Names
Bjarki Ãrmannsson, Hinrik Hafsteinsson, Atli Jasonarson +1
This paper presents the submission of the Ãrni Magnússon Institute's team to the WMT24 test suite subtask, focusing on idiomatic expressions and proper names for the English->Ice…
Cogs in a Machine, Doing What They're Meant to Do -- The AMI Submission to the WMT24 General Translation Task
Atli Jasonarson, Hinrik Hafsteinsson, Bjarki Ãrmannsson +1
This paper presents the submission of the Ãrni Magnusson Institute's team to the WMT24 General translation task. We work on the English->Icelandic translation direction. Our syste…