activity
20202026
most citedNoCoLA: The Norwegian Corpus of Linguistic Acceptability

4 citations · 14 across the 16 of their papers we have counts for

collaborators
Showing 2023Show all

6 papers · 1 filter

cs.CL20232 cited

Not all layers are equally as important: Every Layer Counts BERT

Lucas Georges Gabriel Charpentier, David Samuel

This paper introduces a novel modification of the transformer architecture, tailored for the data-efficient pretraining of language models. This aspect is evaluated by participatin…

cs.CL20232 cited

Mean BERTs make erratic language teachers: the effectiveness of latent bootstrapping in low-resource settings

David Samuel

This paper explores the use of latent bootstrapping, an alternative self-supervision technique, for pretraining language models. Unlike the typical practice of using self-supervisi…

cs.CL20234 cited

NoCoLA: The Norwegian Corpus of Linguistic Acceptability

Matias Jentoft, David Samuel

While there has been a surge of large language models for Norwegian in recent years, we lack any tool to evaluate their understanding of grammaticality. We present two new Norwegia…

cs.CL2023

Tokenization with Factorized Subword Encoding

David Samuel, Lilja Øvrelid

In recent years, language models have become increasingly larger and more complex. However, the input representations for these models continue to rely on simple and greedy subword…

cs.CL20233 cited

NorBench -- A Benchmark for Norwegian Language Models

David Samuel, Andrey Kutuzov, Samia Touileb +5

We present NorBench: a streamlined suite of NLP tasks and probes for evaluating Norwegian language models (LMs) on standardized data splits and evaluation metrics. We also introduc…

cs.CL2023

BRENT: Bidirectional Retrieval Enhanced Norwegian Transformer

Lucas Georges Gabriel Charpentier, Sondre Wold, David Samuel +1

Retrieval-based language models are increasingly employed in question-answering tasks. These models search in a corpus of documents for relevant information instead of having all f…