4 citations · 5 across the 3 of their papers we have counts for
3 papers
Text Annotation Handbook: A Practical Guide for Machine Learning Projects
Felix Stollenwerk, Joey Öhman, Danila Petrelli +5
This handbook is a hands-on guide on how to approach text annotation tasks. It provides a gentle introduction to the topic, an overview of theoretical concepts as well as practical…
Annotated Job Ads with Named Entity Recognition
Felix Stollenwerk, Niklas Fastlund, Anna Nyqvist +1
We have trained a named entity recognition (NER) model that screens Swedish job ads for different kinds of useful information (e.g. skills required from a job seeker). It was obtai…
The Nordic Pile: A 1.2TB Nordic Dataset for Language Modeling
Joey Öhman, Severine Verlinden, Ariel Ekgren +5
Pre-training Large Language Models (LLMs) require massive amounts of text data, and the performance of the LLMs typically correlates with the scale and quality of the datasets. Thi…