most citedThe BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

65 citations · 71 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20241 cited

Vid3D: Synthesis of Dynamic 3D Scenes using 2D Video Diffusion

Rishab Parthasarathy, Zachary Ankner, Aaron Gokaslan

A recent frontier in computer vision has been the task of 3D video generation, which consists of generating a time-varying 3D representation of a scene. To generate dynamic 3D scen…

cs.CL20241 cited

Self-Directed Synthetic Dialogues and Revisions Technical Report

Nathan Lambert, Hailey Schoelkopf, Aaron Gokaslan +3

Synthetic data has become an important tool in the fine-tuning of language models to follow instructions and solve complex problems. Nevertheless, the majority of open data to date…

cs.SE20242 cited

On the Standardization of Behavioral Use Clauses and Their Adoption for Responsible Licensing of AI

Daniel McDuff, Tim Korjakow, Scott Cambo +9

Growing concerns over negligent or malicious uses of AI have increased the appetite for tools that help manage the risks of the technology. In 2018, licenses with behaviorial-use c…

cs.CV20232 cited

CommonCanvas: An Open Diffusion Model Trained with Creative-Commons Images

Aaron Gokaslan, A. Feder Cooper, Jasmine Collins +6

We assemble a dataset of Creative-Commons-licensed (CC) images, which we use to train a set of open diffusion models that are qualitatively competitive with Stable Diffusion 2 (SD2…

cs.CL202365 cited

The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

Hugo Laurençon, Lucile Saulnier, Thomas Wang +51

As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…