12 citations · 20 across the 5 of their papers we have counts for
7 papers
Findings of the VarDial Evaluation Campaign 2023
Noëmi Aepli, Çağrı Çöltekin, Rob Van Der Goot +7
This report presents the results of the shared tasks organized as part of the VarDial Evaluation Campaign 2023. The campaign is part of the tenth workshop on Natural Language Proce…
Language Variety Identification with True Labels
Marcos Zampieri, Kai North, Tommi Jauhiainen +4
Language identification is an important first step in many IR and NLP applications. Most publicly available language identification datasets, however, are compiled under the assump…
Comparing Approaches to Dravidian Language Identification
Tommi Jauhiainen, Tharindu Ranasinghe, Marcos Zampieri
This paper describes the submissions by team HWR to the Dravidian Language Identification (DLI) shared task organized at VarDial 2021 workshop. The DLI training set includes 16,674…
Uralic Language Identification (ULI) 2020 shared task dataset and the Wanca 2017 corpus
Tommi Jauhiainen, Heidi Jauhiainen, Niko Partanen +1
This article introduces the Wanca 2017 corpus of texts crawled from the internet from which the sentences in rare Uralic languages for the use of the Uralic Language Identification…
Language Model Adaptation for Language and Dialect Identification of Text
Tommi Jauhiainen, Krister Lindén, Heidi Jauhiainen
This article describes an unsupervised language model adaptation approach that can be used to enhance the performance of language identification methods. The approach is applied to…
Language and Dialect Identification of Cuneiform Texts
Tommi Jauhiainen, Heidi Jauhiainen, Tero Alstola +1
This article introduces a corpus of cuneiform texts from which the dataset for the use of the Cuneiform Language Identification (CLI) 2019 shared task was derived as well as some p…