4 papers
Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
David Samuel, Lilja Ãvrelid, Erik Velldal +1
We propose a post-training method for lower-resource languages that preserves the fluency of language models even when aligned by disfluent reward models. Preference optimization i…
NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark
Vladislav Mikhailov, Tita Enstad, David Samuel +4
This paper introduces NorEval, a new and comprehensive evaluation suite for large-scale standardized benchmarking of Norwegian generative language models (LMs). NorEval consists of…
Small Languages, Big Models: A Study of Continual Training on Languages of Norway
David Samuel, Vladislav Mikhailov, Erik Velldal +4
Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages l…
The Impact of Copyrighted Material on Large Language Models: A Norwegian Perspective
Javier de la Rosa, Vladislav Mikhailov, Lemei Zhang +16
The use of copyrighted materials in training language models raises critical legal and ethical questions. This paper presents a framework for and the results of empirically assessi…