3 papers
cs.CL2026
How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
Kushal Tatariya, Artur Kulmizev, Wessel Poelman +6
Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have…
cs.CL2026
A Dataset for Probing Translationese Preferences in English-to-Swedish Translation
Jenny Kunz, Anja Jarochenko, Marcel Bollmann
Translations often carry traces of the source language, a phenomenon known as translationese. We introduce the first freely available English-to-Swedish dataset contrasting transla…
cs.CL2025
Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation
Kevin Glocker, Kätriin Kukk, Romina Oji +3
Achieving high-performing language models which include medium- and lower-resource languages remains a challenge. Massively multilingual models still underperform compared to langu…