most citedThe Nordic Pile: A 1.2TB Nordic Dataset for Language Modeling

4 citations · 8 across the 2 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2023★ 4 cited

GPT-SW3: An Autoregressive Language Model for the Nordic Languages

Ariel Ekgren, Amaru Cuba Gyllensten, Felix Stollenwerk +7

This paper details the process of developing the first native large generative language model for the Nordic languages, GPT-SW3. We cover all parts of the development process, from…

cs.CL2023★ 4 cited

The Nordic Pile: A 1.2TB Nordic Dataset for Language Modeling

Joey Öhman, Severine Verlinden, Ariel Ekgren +5

Pre-training Large Language Models (LLMs) require massive amounts of text data, and the performance of the LLMs typically correlates with the scale and quality of the datasets. Thi…

cs.CL2018

Measuring Issue Ownership using Word Embeddings

Amaru Cuba Gyllensten, Magnus Sahlgren

Sentiment and topic analysis are common methods used for social media monitoring. Essentially, these methods answers questions such as, "what is being talked about, regarding X", a…

cs.CL2018

R-grams: Unsupervised Learning of Semantic Units in Natural Language

Ariel Ekgren, Amaru Cuba Gyllensten, Magnus Sahlgren

This paper investigates data-driven segmentation using Re-Pair or Byte Pair Encoding-techniques. In contrast to previous work which has primarily been focused on subword units for…

cs.CL2018

Distributional Term Set Expansion

Amaru Cuba Gyllensten, Magnus Sahlgren

This paper is a short empirical study of the performance of centrality and classification based iterative term set expansion methods for distributional semantic models. Iterative t…