collaborators

6 papers

cs.CL2025

An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)

Laurie Burchell, Ona de Gibert, Nikolay Arefyev +32

Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In th…

cs.CL2025

Multi-label Scandinavian Language Identification (SLIDE)

Mariia Fedorova, Jonas Sebulon Frydenberg, Victoria Handford +6

Identifying closely related languages at sentence level is difficult, in particular because it is often impossible to assign a sentence to a single language. In this paper, we focu…

cs.CL2025

Mixed Feelings: Cross-Domain Sentiment Classification of Patient Feedback

Egil Rønningstad, Lilja Charlotte Storset, Petter Mæhlum +2

Sentiment analysis of patient feedback from the public health domain can aid decision makers in evaluating the provided services. The current paper focuses on free-text comments in…

cs.CL2025

The Impact of Copyrighted Material on Large Language Models: A Norwegian Perspective

Javier de la Rosa, Vladislav Mikhailov, Lemei Zhang +16

The use of copyrighted materials in training language models raises critical legal and ethical questions. This paper presents a framework for and the results of empirically assessi…

cs.CL2025

A Collection of Question Answering Datasets for Norwegian

Vladislav Mikhailov, Petter Mæhlum, Victoria Ovedie Chruickshank Langø +2

This paper introduces a new suite of question answering datasets for Norwegian; NorOpenBookQA, NorCommonSenseQA, NorTruthfulQA, and NRK-Quiz-QA. The data covers a wide range of ski…

cs.CL2024

It's Difficult to be Neutral -- Human and LLM-based Sentiment Annotation of Patient Comments

Petter Mæhlum, David Samuel, Rebecka Maria Norman +4

Sentiment analysis is an important tool for aggregating patient voices, in order to provide targeted improvements in healthcare services. A prerequisite for this is the availabilit…