2 papers
cs.CL2025
BiaSWE: An Expert Annotated Dataset for Misogyny Detection in Swedish
Kätriin Kukk, Danila Petrelli, Judit Casademont +3
In this study, we introduce the process for creating BiaSWE, an expert-annotated dataset tailored for misogyny detection in the Swedish language. To address the cultural and lingui…
cs.CL2024
SWEb: A Large Web Dataset for the Scandinavian Languages
Tobias Norlund, Tim Isbister, Amaru Cuba Gyllensten +4
This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the…