most citedThe BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

65 citations · 130 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2023

GAIA Search: Hugging Face and Pyserini Interoperability for NLP Training Data Exploration

Aleksandra Piktus, Odunayo Ogundepo, Christopher Akiki +6

Noticing the urgent need to provide tools for fast and user-friendly qualitative analysis of large-scale textual corpora of the modern NLP, we propose to turn to the mature and wel…

cs.CL202365 cited

The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

Hugo Laurençon, Lucile Saulnier, Thomas Wang +51

As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…

cs.CL2023

The ROOTS Search Tool: Data Transparency for LLMs

Aleksandra Piktus, Christopher Akiki, Paulo Villegas +5

ROOTS is a 1.6TB multilingual text corpus developed for the training of BLOOM, currently the largest language model explicitly accompanied by commensurate data governance efforts.…

cs.SE202354 cited

SantaCoder: don't reach for the stars!

Loubna Ben Allal, Raymond Li, Denis Kocetkov +38

The BigCode project is an open-scientific collaboration working on the responsible development of large language models for code. This tech report describes the progress of the col…

cs.CY202311 cited

Towards Openness Beyond Open Access: User Journeys through 3 Open AI Collaboratives

Jennifer Ding, Christopher Akiki, Yacine Jernite +2

Open Artificial Intelligence (Open source AI) collaboratives offer alternative pathways for how AI can be developed beyond well-resourced technology companies and who can be a part…