3 papers
cs.CL2026
OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report
Mariia Fedorova, Nikolay Arefyev, Maja Buljan +4
Language identification (LID) is an essential step in building high-quality multilingual datasets from web data. Existing LID tools (such as OpenLID or GlotLID) often struggle to i…
cs.CL2025
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
Laurie Burchell, Ona de Gibert, Nikolay Arefyev +32
Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In th…
cs.CL2025
Multi-label Scandinavian Language Identification (SLIDE)
Mariia Fedorova, Jonas Sebulon Frydenberg, Victoria Handford +6
Identifying closely related languages at sentence level is difficult, in particular because it is often impossible to assign a sentence to a single language. In this paper, we focu…