65 citations · 117 across the 7 of their papers we have counts for
5 papers · 1 filter
CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
Vipul Gupta, Pranav Narayanan Venkit, Hugo Laurençon +2
As language models (LMs) become increasingly powerful and widely used, it is important to quantify them for sociodemographic bias with potential for harm. Prior measures of bias ar…
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…
The ROOTS Search Tool: Data Transparency for LLMs
Aleksandra Piktus, Christopher Akiki, Paulo Villegas +5
ROOTS is a 1.6TB multilingual text corpus developed for the training of BLOOM, currently the largest language model explicitly accompanied by commensurate data governance efforts.…
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
BigScience Workshop, :, Teven Le Scao +391
Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to wi…
DP-Parse: Finding Word Boundaries from Raw Speech with an Instance Lexicon
Robin Algayres, Tristan Ricoul, Julien Karadayi +5
Finding word boundaries in continuous speech is challenging as there is little or no equivalent of a 'space' delimiter between words. Popular Bayesian non-parametric models for tex…