papers

Publications (11)

cs.CL2022

A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation

David Ifeoluwa Adelani, Jesujoba Oluwadara Alabi, Angela Fan +42

Recent advances in the pre-training of language models leverage large-scale datasets to create multilingual models. However, low-resource languages are mostly left out in these dat…

cs.CY2019

Keyword Spotter Model for Crop Pest and Disease Monitoring from Community Radio Data

Benjamin Akera, Joyce Nakatumba-Nabende, Jonathan Mukiibi +4

In societies with well developed internet infrastructure, social media is the leading medium of communication for various social issues especially for breaking news situations. In…

cs.CL2023

MasakhaNEWS: News Topic Classification for African languages

David Ifeoluwa Adelani, Marek Masiak, Israel Abebe Azime +62

African languages are severely under-represented in NLP research due to lack of datasets covering several NLP tasks. While there are individual language specific datasets that are…

cs.CL2023

MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African Languages

Cheikh M. Bamba Dione, David Adelani, Peter Nabende +41

In this paper, we present MasakhaPOS, the largest part-of-speech (POS) dataset for 20 typologically diverse African languages. We discuss the challenges in annotating POS for these…

eess.AS2022

BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus

Josh Meyer, David Ifeoluwa Adelani, Edresson Casanova +16

BibleTTS is a large, high-quality, open speech dataset for ten languages spoken in Sub-Saharan Africa. The corpus contains up to 86 hours of aligned, studio quality 48kHz single sp…

cs.CL2025

IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models

David Ifeoluwa Adelani, Jessica Ojo, Israel Abebe Azime +24

Despite the widespread adoption of Large language models (LLMs), their remarkable capabilities remain limited to a few high-resource languages. Additionally, many low-resource lang…