activity
20222025
most citedThe Effectiveness of Masked Language Modeling and Adapters for Factual Knowledge Injection

1 citations · 1 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2025

Systematic Generalization in Language Models Scales with Information Entropy

Sondre Wold, Lucas Georges Gabriel Charpentier, Étienne Simon

Systematic generalization remains challenging for current language models, which are known to be both sensitive to semantically similar permutations of the input and to struggle wi…

cs.CL2024

Compositional Generalization with Grounded Language Models

Sondre Wold, Étienne Simon, Lucas Georges Gabriel Charpentier +3

Grounded language models use external sources of information, such as knowledge graphs, to meet some of the general challenges associated with pre-training. By extending previous w…

cs.CL2024

More Room for Language: Investigating the Effect of Retrieval on Language Models

David Samuel, Lucas Georges Gabriel Charpentier, Sondre Wold

Retrieval-augmented language models pose a promising alternative to standard language modeling. During pretraining, these models search in a corpus of documents for contextually re…

cs.CL2024

Estimating Lexical Complexity from Document-Level Distributions

Sondre Wold, Petter Mæhlum, Oddbjørn Hove

Existing methods for complexity estimation are typically developed for entire documents. This limitation in scope makes them inapplicable for shorter pieces of text, such as health…

cs.CL2023

Text-To-KG Alignment: Comparing Current Methods on Classification Tasks

Sondre Wold, Lilja Øvrelid, Erik Velldal

In contrast to large text corpora, knowledge graphs (KG) provide dense and structured representations of factual information. This makes them attractive for systems that supplement…

cs.CL2023

NorQuAD: Norwegian Question Answering Dataset

Sardana Ivanova, Fredrik Aas Andreassen, Matias Jentoft +2

In this paper we present NorQuAD: the first Norwegian question answering dataset for machine reading comprehension. The dataset consists of 4,752 manually created question-answer p…