4 citations · 4 across the 2 of their papers we have counts for
4 papers
SWEb: A Large Web Dataset for the Scandinavian Languages
Tobias Norlund, Tim Isbister, Amaru Cuba Gyllensten +4
This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the…
GPT-SW3: An Autoregressive Language Model for the Nordic Languages
Ariel Ekgren, Amaru Cuba Gyllensten, Felix Stollenwerk +7
This paper details the process of developing the first native large generative language model for the Nordic languages, GPT-SW3. We cover all parts of the development process, from…
The Nordic Pile: A 1.2TB Nordic Dataset for Language Modeling
Joey Öhman, Severine Verlinden, Ariel Ekgren +5
Pre-training Large Language Models (LLMs) require massive amounts of text data, and the performance of the LLMs typically correlates with the scale and quality of the datasets. Thi…
R-grams: Unsupervised Learning of Semantic Units in Natural Language
Ariel Ekgren, Amaru Cuba Gyllensten, Magnus Sahlgren
This paper investigates data-driven segmentation using Re-Pair or Byte Pair Encoding-techniques. In contrast to previous work which has primarily been focused on subword units for…