3 citations · 5 across the 4 of their papers we have counts for
6 papers · 1 filter
Small Languages, Big Models: A Study of Continual Training on Languages of Norway
David Samuel, Vladislav Mikhailov, Erik Velldal +4
Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages l…
More Room for Language: Investigating the Effect of Retrieval on Language Models
David Samuel, Lucas Georges Gabriel Charpentier, Sondre Wold
Retrieval-augmented language models pose a promising alternative to standard language modeling. During pretraining, these models search in a corpus of documents for contextually re…
Not all layers are equally as important: Every Layer Counts BERT
Lucas Georges Gabriel Charpentier, David Samuel
This paper introduces a novel modification of the transformer architecture, tailored for the data-efficient pretraining of language models. This aspect is evaluated by participatin…
NorBench -- A Benchmark for Norwegian Language Models
David Samuel, Andrey Kutuzov, Samia Touileb +5
We present NorBench: a streamlined suite of NLP tasks and probes for evaluating Norwegian language models (LMs) on standardized data splits and evaluation metrics. We also introduc…
BRENT: Bidirectional Retrieval Enhanced Norwegian Transformer
Lucas Georges Gabriel Charpentier, Sondre Wold, David Samuel +1
Retrieval-based language models are increasingly employed in question-answering tasks. These models search in a corpus of documents for relevant information instead of having all f…
Trained on 100 million words and still in shape: BERT meets British National Corpus
David Samuel, Andrey Kutuzov, Lilja Øvrelid +1
While modern masked language models (LMs) are trained on ever larger corpora, we here explore the effects of down-scaling training to a modestly-sized but representative, well-bala…