most citedThe BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

65 citations · 72 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL20231 cited

MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African Languages

Cheikh M. Bamba Dione, David Adelani, Peter Nabende +41

In this paper, we present MasakhaPOS, the largest part-of-speech (POS) dataset for 20 typologically diverse African languages. We discuss the challenges in annotating POS for these…

cs.CL20234 cited

SemEval-2023 Task 12: Sentiment Analysis for African Languages (AfriSenti-SemEval)

Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Seid Muhie Yimam +7

We present the first Africentric SemEval Shared task, Sentiment Analysis for African Languages (AfriSenti-SemEval) - The dataset is available at https://github.com/afrisenti-semeva…

cs.CL2023

KÚ <MASK>: Integrating Yorùbá cultural greetings into machine translation

Idris Akinade, Jesujoba Alabi, David Adelani +2

This paper investigates the performance of massively multilingual neural machine translation (NMT) systems in translating Yorùbá greetings ( kú [MASK]), which are a bi…

cs.CL20232 cited

MphayaNER: Named Entity Recognition for Tshivenda

Rendani Mbuvha, David I. Adelani, Tendani Mutavhatsindi +7

Named Entity Recognition (NER) plays a vital role in various Natural Language Processing tasks such as information retrieval, text classification, and question answering. However,…

cs.CL202365 cited

The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

Hugo Laurençon, Lucile Saulnier, Thomas Wang +51

As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…

eess.AS2022

BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus

Josh Meyer, David Ifeoluwa Adelani, Edresson Casanova +16

BibleTTS is a large, high-quality, open speech dataset for ten languages spoken in Sub-Saharan Africa. The corpus contains up to 86 hours of aligned, studio quality 48kHz single sp…