4 citations · 4 across the 2 of their papers we have counts for
3 papers
SWEb: A Large Web Dataset for the Scandinavian Languages
Tobias Norlund, Tim Isbister, Amaru Cuba Gyllensten +4
This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the…
Should we Stop Training More Monolingual Models, and Simply Use Machine Translation Instead?
Tim Isbister, Fredrik Carlsson, Magnus Sahlgren
Most work in NLP makes the assumption that it is desirable to develop solutions in the native language in question. There is consequently a strong trend towards building native lan…
Federated Word2Vec: Leveraging Federated Learning to Encourage Collaborative Representation Learning
Daniel Garcia Bernal, Lodovico Giaretta, Sarunas Girdzijauskas +1
Large scale contextual representation models have significantly advanced NLP in recent years, understanding the semantics of text to a degree never seen before. However, they need…