65 citations · 66 across the 3 of their papers we have counts for
3 papers
AfroDigits: A Community-Driven Spoken Digit Dataset for African Languages
Chris Chinenye Emezue, Sanchit Gandhi, Lewis Tunstall +10
The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address thi…
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…
Training Transformers Together
Alexander Borzunov, Max Ryabinin, Tim Dettmers +5
The infrastructure necessary for training state-of-the-art models is becoming overly expensive, which makes training such models affordable only to large corporations and instituti…