65 citations · 130 across the 5 of their papers we have counts for
5 papers
GAIA Search: Hugging Face and Pyserini Interoperability for NLP Training Data Exploration
Aleksandra Piktus, Odunayo Ogundepo, Christopher Akiki +6
Noticing the urgent need to provide tools for fast and user-friendly qualitative analysis of large-scale textual corpora of the modern NLP, we propose to turn to the mature and wel…
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…
The ROOTS Search Tool: Data Transparency for LLMs
Aleksandra Piktus, Christopher Akiki, Paulo Villegas +5
ROOTS is a 1.6TB multilingual text corpus developed for the training of BLOOM, currently the largest language model explicitly accompanied by commensurate data governance efforts.…
SantaCoder: don't reach for the stars!
Loubna Ben Allal, Raymond Li, Denis Kocetkov +38
The BigCode project is an open-scientific collaboration working on the responsible development of large language models for code. This tech report describes the progress of the col…
Towards Openness Beyond Open Access: User Journeys through 3 Open AI Collaboratives
Jennifer Ding, Christopher Akiki, Yacine Jernite +2
Open Artificial Intelligence (Open source AI) collaboratives offer alternative pathways for how AI can be developed beyond well-resourced technology companies and who can be a part…