3 papers
cs.CL2020
Uralic Language Identification (ULI) 2020 shared task dataset and the Wanca 2017 corpus
Tommi Jauhiainen, Heidi Jauhiainen, Niko Partanen +1
This article introduces the Wanca 2017 corpus of texts crawled from the internet from which the sentences in rare Uralic languages for the use of the Uralic Language Identification…
cs.CL2019
Language Model Adaptation for Language and Dialect Identification of Text
Tommi Jauhiainen, Krister Lindén, Heidi Jauhiainen
This article describes an unsupervised language model adaptation approach that can be used to enhance the performance of language identification methods. The approach is applied to…
cs.CL2019
Language and Dialect Identification of Cuneiform Texts
Tommi Jauhiainen, Heidi Jauhiainen, Tero Alstola +1
This article introduces a corpus of cuneiform texts from which the dataset for the use of the Cuneiform Language Identification (CLI) 2019 shared task was derived as well as some p…