65 citations · 303 across the 22 of their papers we have counts for
13 papers · 1 filter
INTIMA: A Benchmark for Human-AI Companionship Behavior
Lucie-Aimée Kaffee, Giada Pistilli, Yacine Jernite
AI companionship, where users develop emotional bonds with AI systems, has emerged as a significant pattern with positive but also concerning implications. We introduce Interaction…
CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models
Giada Pistilli, Alina Leidinger, Yacine Jernite +3
This paper introduces the "CIVICS: Culturally-Informed & Values-Inclusive Corpus for Societal impacts" dataset, designed to evaluate the social and cultural variation of Large Lang…
StarCoder: may the source be with you!
Raymond Li, Loubna Ben Allal, Yangtian Zi +64
The BigCode community, an open-scientific collaboration working on the responsible development of Large Language Models for Code (Code LLMs), introduces StarCoder and StarCoderBase…
AfroDigits: A Community-Driven Spoken Digit Dataset for African Languages
Chris Chinenye Emezue, Sanchit Gandhi, Lewis Tunstall +10
The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address thi…
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…
The Stack: 3 TB of permissively licensed source code
Denis Kocetkov, Raymond Li, Loubna Ben Allal +10
Large Language Models (LLMs) play an ever-increasing role in the field of Artificial Intelligence (AI)--not only for natural language processing but also for code understanding and…