4 citations · 14 across the 16 of their papers we have counts for
6 papers · 1 filter
Not all layers are equally as important: Every Layer Counts BERT
Lucas Georges Gabriel Charpentier, David Samuel
This paper introduces a novel modification of the transformer architecture, tailored for the data-efficient pretraining of language models. This aspect is evaluated by participatin…
Mean BERTs make erratic language teachers: the effectiveness of latent bootstrapping in low-resource settings
David Samuel
This paper explores the use of latent bootstrapping, an alternative self-supervision technique, for pretraining language models. Unlike the typical practice of using self-supervisi…
NoCoLA: The Norwegian Corpus of Linguistic Acceptability
Matias Jentoft, David Samuel
While there has been a surge of large language models for Norwegian in recent years, we lack any tool to evaluate their understanding of grammaticality. We present two new Norwegia…
Tokenization with Factorized Subword Encoding
David Samuel, Lilja Øvrelid
In recent years, language models have become increasingly larger and more complex. However, the input representations for these models continue to rely on simple and greedy subword…
NorBench -- A Benchmark for Norwegian Language Models
David Samuel, Andrey Kutuzov, Samia Touileb +5
We present NorBench: a streamlined suite of NLP tasks and probes for evaluating Norwegian language models (LMs) on standardized data splits and evaluation metrics. We also introduc…
BRENT: Bidirectional Retrieval Enhanced Norwegian Transformer
Lucas Georges Gabriel Charpentier, Sondre Wold, David Samuel +1
Retrieval-based language models are increasingly employed in question-answering tasks. These models search in a corpus of documents for relevant information instead of having all f…