activity
20212023
most citedBloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks

2 citations · 7 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2023

The eBible Corpus: Data and Model Benchmarks for Bible Translation for Low-Resource Languages

Vesa Akerman, David Baines, Damien Daspit +7

Efficiently and accurately translating a corpus into a low-resource language remains a challenge, regardless of the strategies employed, whether manual, automated, or a combination…

cs.CL2023★ 2 cited

Adapting to the Low-Resource Double-Bind: Investigating Low-Compute Methods on Low-Resource African Languages

Colin Leong, Herumb Shandilya, Bonaventure F. P. Dossou +10

Many natural language processing (NLP) tasks make use of massively pre-trained language models, which are computationally expensive. However, access to high computational resources…

cs.CL2022

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

BigScience Workshop, :, Teven Le Scao +391

Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to wi…

cs.CL2022★ 2 cited

Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks

Colin Leong, Joshua Nemecek, Jacob Mansdorfer +3

We present Bloom Library, a linguistically diverse set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/re…

cs.CL2022★ 1 cited

A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation

David Ifeoluwa Adelani, Jesujoba Oluwadara Alabi, Angela Fan +42

Recent advances in the pre-training of language models leverage large-scale datasets to create multilingual models. However, low-resource languages are mostly left out in these dat…

cs.CL2022★ 2 cited

Documenting Geographically and Contextually Diverse Data Sources: The BigScience Catalogue of Language Data and Resources

Angelina McMillan-Major, Zaid Alyafeai, Stella Biderman +15

In recent years, large-scale data collection efforts have prioritized the amount of data collected in order to improve the modeling capabilities of large language models. This prio…