collaborators

6 papers

cs.LG2026

SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models

Hoda Fakharzadehjahromy, Emil Wiman, Andreas Bueff +2

Preference optimisation has proven effective for improving large language models but typically relies on costly human preference annotations. Extending these methods to morphologic…

cs.CY2026

Curated retrieval versus open web search in public AI information services: a coverage-trust trade-off

Hafsteinn Einarsson, Hafsteinn Birgir Einarsson, Jón Gunnar Ólafsson +1

Public institutions increasingly use large language models (LLMs) to answer citizens' questions, often pairing a curated knowledge base with live web search, yet whether the source…

cs.CL2026

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…

cs.AI2025

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Hafsteinn Einarsson

As Large Language Models (LLMs) increasingly power autonomous agents in robotics and embodied AI, understanding their spatial reasoning capabilities becomes crucial for ensuring re…

cs.CL2025

Hotter and Colder: A New Approach to Annotating Sentiment, Emotions, and Bias in Icelandic Blog Comments

Steinunn Rut Friðriksdóttir, Dan Saattrup Nielsen, Hafsteinn Einarsson

This paper presents Hotter and Colder, a dataset designed to analyze various types of online behavior in Icelandic blog comments. Building on previous work, we used GPT-4o mini to…

cs.CL2025

FoQA: A Faroese Question-Answering Dataset

Annika Simonsen, Dan Saattrup Nielsen, Hafsteinn Einarsson

We present FoQA, a Faroese extractive question-answering (QA) dataset with 2,000 samples, created using a semi-automated approach combining Large Language Models (LLMs) and human v…