6 citations · 7 across the 3 of their papers we have counts for
3 papers · 1 filter
FastSpell: the LangId Magic Spell
Marta Bañón, Jaume Zaragoza-Bernabeu, Gema Ramírez-Sánchez +1
Language identification is a crucial component in the automated production of language resources, particularly in multilingual and big data contexts. However, commonly used languag…
A New Massive Multilingual Dataset for High-Performance Language Technologies
Ona de Gibert, Graeme Nail, Nikolay Arefyev +10
We present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from…
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
Rik van Noord, Taja Kuzman, Peter Rupnik +4
Large, curated, web-crawled corpora play a vital role in training language models (LMs). They form the lion's share of the training data in virtually all recent LMs, such as the we…