3 papers
cs.CL2026
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
Maximilian Idahl, Jörg Tiedemann, Sampo Pyysalo +19
Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approxi…
cs.CL2025
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
Laurie Burchell, Ona de Gibert, Nikolay Arefyev +32
Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In th…
cs.CL2024
FastSpell: the LangId Magic Spell
Marta Bañón, Jaume Zaragoza-Bernabeu, Gema RamÃrez-Sánchez +1
Language identification is a crucial component in the automated production of language resources, particularly in multilingual and big data contexts. However, commonly used languag…