4 papers
A Finnish News Corpus for Named Entity Recognition
Teemu Ruokolainen, Pekka Kauppinen, Miikka Silfverberg +1
We present a corpus of Finnish news articles with a manually prepared named entity annotation. The corpus consists of 953 articles (193,742 word tokens) with six named entity class…
Language Model Adaptation for Language and Dialect Identification of Text
Tommi Jauhiainen, Krister Lindén, Heidi Jauhiainen
This article describes an unsupervised language model adaptation approach that can be used to enhance the performance of language identification methods. The approach is applied to…
Language and Dialect Identification of Cuneiform Texts
Tommi Jauhiainen, Heidi Jauhiainen, Tero Alstola +1
This article introduces a corpus of cuneiform texts from which the dataset for the use of the Cuneiform Language Identification (CLI) 2019 shared task was derived as well as some p…
Automatic Language Identification in Texts: A Survey
Tommi Jauhiainen, Marco Lui, Marcos Zampieri +2
Language identification (LI) is the problem of determining the natural language that a document or part thereof is written in. Automatic LI has been extensively researched for over…