6 papers
SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models
Hoda Fakharzadehjahromy, Emil Wiman, Andreas Bueff +2
Preference optimisation has proven effective for improving large language models but typically relies on costly human preference annotations. Extending these methods to morphologic…
Curated retrieval versus open web search in public AI information services: a coverage-trust trade-off
Hafsteinn Einarsson, Hafsteinn Birgir Einarsson, Jón Gunnar Ãlafsson +1
Public institutions increasingly use large language models (LLMs) to answer citizens' questions, often pairing a curated knowledge base with live web search, yet whether the source…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
Hafsteinn Einarsson
As Large Language Models (LLMs) increasingly power autonomous agents in robotics and embodied AI, understanding their spatial reasoning capabilities becomes crucial for ensuring re…
Hotter and Colder: A New Approach to Annotating Sentiment, Emotions, and Bias in Icelandic Blog Comments
Steinunn Rut Friðriksdóttir, Dan Saattrup Nielsen, Hafsteinn Einarsson
This paper presents Hotter and Colder, a dataset designed to analyze various types of online behavior in Icelandic blog comments. Building on previous work, we used GPT-4o mini to…
FoQA: A Faroese Question-Answering Dataset
Annika Simonsen, Dan Saattrup Nielsen, Hafsteinn Einarsson
We present FoQA, a Faroese extractive question-answering (QA) dataset with 2,000 samples, created using a semi-automated approach combining Large Language Models (LLMs) and human v…