2 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CL2022★ 2 cited
Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks
Colin Leong, Joshua Nemecek, Jacob Mansdorfer +3
We present Bloom Library, a linguistically diverse set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/re…
cs.CL2022★ 2 cited
Documenting Geographically and Contextually Diverse Data Sources: The BigScience Catalogue of Language Data and Resources
Angelina McMillan-Major, Zaid Alyafeai, Stella Biderman +15
In recent years, large-scale data collection efforts have prioritized the amount of data collected in order to improve the modeling capabilities of large language models. This prio…