activity
20182021
most citedData Readiness for Natural Language Processing

1 citations · 1 across the 3 of their papers we have counts for

collaborators

8 papers

cs.CL2021

We Need to Talk About Data: The Importance of Data Readiness in Natural Language Processing

Fredrik Olsson, Magnus Sahlgren

In this paper, we identify the state of data as being an important reason for failure in applied Natural Language Processing (NLP) projects. We argue that there is a gap between ac…

cs.CL2021

The Singleton Fallacy: Why Current Critiques of Language Models Miss the Point

Magnus Sahlgren, Fredrik Carlsson

This paper discusses the current critique against neural network-based Natural Language Understanding (NLU) solutions known as language models. We argue that much of the current de…

cs.CY20201 cited

Data Readiness for Natural Language Processing

Fredrik Olsson, Magnus Sahlgren

This document concerns data readiness in the context of machine learning and Natural Language Processing. It describes how an organization may proceed to identify, make available,…

cs.CL2020

Why Not Simply Translate? A First Swedish Evaluation Benchmark for Semantic Similarity

Tim Isbister, Magnus Sahlgren

This paper presents the first Swedish evaluation benchmark for textual semantic similarity. The benchmark is compiled by simply running the English STS-B dataset through the Google…

cs.CL2018

Measuring Issue Ownership using Word Embeddings

Amaru Cuba Gyllensten, Magnus Sahlgren

Sentiment and topic analysis are common methods used for social media monitoring. Essentially, these methods answers questions such as, "what is being talked about, regarding X", a…

cs.CL2018

R-grams: Unsupervised Learning of Semantic Units in Natural Language

Ariel Ekgren, Amaru Cuba Gyllensten, Magnus Sahlgren

This paper investigates data-driven segmentation using Re-Pair or Byte Pair Encoding-techniques. In contrast to previous work which has primarily been focused on subword units for…