3 papers
cs.CL2026
How Should We Model the Probability of a Language?
Rasul Dent, Pedro Ortiz Suarez, Thibault Clérice +1
Of the over 7,000 languages spoken in the world, commercial language identification (LID) systems only reliably identify a few hundred in written form. Research-grade systems exten…
cs.CL2025
Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
Ekaterina Artemova, Laurie Burchell, Daryna Dementieva +3
This tutorial (https://tum-nlp.github.io/low-resource-tutorial) is designed for NLP practitioners, researchers, and developers working with multilingual and low-resource languages…
cs.CL2025
KréyoLID From Language Identification Towards Language Mining
Rasul Dent, Pedro Ortiz Suarez, Thibault Clérice +1
Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may b…