4 citations · 8 across the 3 of their papers we have counts for
1 paper · 2 filters
Tobias Norlund, Tim Isbister, Amaru Cuba Gyllensten +4
This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the…