3 papers
cs.CL2025
The Impact of Copyrighted Material on Large Language Models: A Norwegian Perspective
Javier de la Rosa, Vladislav Mikhailov, Lemei Zhang +16
The use of copyrighted materials in training language models raises critical legal and ethical questions. This paper presents a framework for and the results of empirically assessi…
cs.CL2025
A Collection of Question Answering Datasets for Norwegian
Vladislav Mikhailov, Petter Mæhlum, Victoria Ovedie Chruickshank Langø +2
This paper introduces a new suite of question answering datasets for Norwegian; NorOpenBookQA, NorCommonSenseQA, NorTruthfulQA, and NRK-Quiz-QA. The data covers a wide range of ski…
cs.CL2025
Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles
Samia Touileb, Vladislav Mikhailov, Marie Kroka +2
We introduce a dataset of high-quality human-authored summaries of news articles in Norwegian. The dataset is intended for benchmarking the abstractive summarisation capabilities o…