65 citations · 68 across the 2 of their papers we have counts for
2 papers
cs.CL2023★ 65 cited
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…
cs.CL2022★ 3 cited
Masader Plus: A New Interface for Exploring +500 Arabic NLP Datasets
Yousef Altaher, Ali Fadel, Mazen Alotaibi +18
Masader (Alyafeai et al., 2021) created a metadata structure to be used for cataloguing Arabic NLP datasets. However, developing an easy way to explore such a catalogue is a challe…