6 papers
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
Laurie Burchell, Ona de Gibert, Nikolay Arefyev +32
Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In th…
Multi-label Scandinavian Language Identification (SLIDE)
Mariia Fedorova, Jonas Sebulon Frydenberg, Victoria Handford +6
Identifying closely related languages at sentence level is difficult, in particular because it is often impossible to assign a sentence to a single language. In this paper, we focu…
Mixed Feelings: Cross-Domain Sentiment Classification of Patient Feedback
Egil Rønningstad, Lilja Charlotte Storset, Petter Mæhlum +2
Sentiment analysis of patient feedback from the public health domain can aid decision makers in evaluating the provided services. The current paper focuses on free-text comments in…
The Impact of Copyrighted Material on Large Language Models: A Norwegian Perspective
Javier de la Rosa, Vladislav Mikhailov, Lemei Zhang +16
The use of copyrighted materials in training language models raises critical legal and ethical questions. This paper presents a framework for and the results of empirically assessi…
A Collection of Question Answering Datasets for Norwegian
Vladislav Mikhailov, Petter Mæhlum, Victoria Ovedie Chruickshank Langø +2
This paper introduces a new suite of question answering datasets for Norwegian; NorOpenBookQA, NorCommonSenseQA, NorTruthfulQA, and NRK-Quiz-QA. The data covers a wide range of ski…
It's Difficult to be Neutral -- Human and LLM-based Sentiment Annotation of Patient Comments
Petter Mæhlum, David Samuel, Rebecka Maria Norman +4
Sentiment analysis is an important tool for aggregating patient voices, in order to provide targeted improvements in healthcare services. A prerequisite for this is the availabilit…