6 citations · 7 across the 4 of their papers we have counts for
2 papers
cs.CL2025
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
Laurie Burchell, Ona de Gibert, Nikolay Arefyev +32
Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In th…
cs.CL2023★ 1 cited
OpusCleaner and OpusTrainer, open source toolkits for training Machine Translation and Large language models
Nikolay Bogoychev, Jelmer van der Linde, Graeme Nail +7
Developing high quality machine translation systems is a labour intensive, challenging and confusing process for newcomers to the field. We present a pair of tools OpusCleaner and…