4 papers
OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report
Mariia Fedorova, Nikolay Arefyev, Maja Buljan +4
Language identification (LID) is an essential step in building high-quality multilingual datasets from web data. Existing LID tools (such as OpenLID or GlotLID) often struggle to i…
Language Ideologies in a Multilingual Society: An LLM-based Analysis of Luxembourgish News Comments
Emilia Milano, Alistair Plum, Yves Scherrer +1
Detecting language ideologies is a valuable yet complex task for understanding how identities are constructed through discourse. In Luxembourg's multicultural and multilingual soci…
Explaining novel senses using definition generation with open language models
Mariia Fedorova, Andrey Kutuzov, Francesco Periti +1
We apply definition generators based on open-weights large language models to the task of creating explanations of novel senses, taking target word usages as an input. To this end,…
Multi-label Scandinavian Language Identification (SLIDE)
Mariia Fedorova, Jonas Sebulon Frydenberg, Victoria Handford +6
Identifying closely related languages at sentence level is difficult, in particular because it is often impossible to assign a sentence to a single language. In this paper, we focu…