65 citations · 66 across the 2 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2023★ 65 cited
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…
cs.CL2021★ 1 cited
Models in the Loop: Aiding Crowdworkers with Generative Annotation Assistants
Max Bartolo, Tristan Thrush, Sebastian Riedel +3
In Dynamic Adversarial Data Collection (DADC), human annotators are tasked with finding examples that models struggle to predict correctly. Models trained on DADC-collected trainin…