3 papers
cs.CL2022
Lahjoita puhetta -- a large-scale corpus of spoken Finnish with some benchmarks
Anssi Moisio, Dejan Porjazovski, Aku Rouhe +5
The Donate Speech campaign has so far succeeded in gathering approximately 3600 hours of ordinary, colloquial Finnish speech into the Lahjoita puhetta (Donate Speech) corpus. The c…
cs.CL2020
Uralic Language Identification (ULI) 2020 shared task dataset and the Wanca 2017 corpus
Tommi Jauhiainen, Heidi Jauhiainen, Niko Partanen +1
This article introduces the Wanca 2017 corpus of texts crawled from the internet from which the sentences in rare Uralic languages for the use of the Uralic Language Identification…
cs.CL2020
The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual Europe
Georg Rehm, Katrin Marheinecke, Stefanie Hegele +44
Multilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cr…