52 citations · 92 across the 13 of their papers we have counts for
19 papers · 1 filter
AfriNames: Most ASR models "butcher" African Names
Tobi Olatunji, Tejumade Afonja, Bonaventure F. P. Dossou +4
Useful conversational agents must accurately capture named entities to minimize error for downstream tasks, for example, asking a voice assistant to play a track from a certain art…
MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African Languages
Cheikh M. Bamba Dione, David Adelani, Peter Nabende +41
In this paper, we present MasakhaPOS, the largest part-of-speech (POS) dataset for 20 typologically diverse African languages. We discuss the challenges in annotating POS for these…
AfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages
Odunayo Ogundepo, Tajuddeen R. Gwadabe, Clara E. Rivera +49
African languages have far less in-language content available digitally, making it challenging for question answering systems to satisfy the information needs of users. Cross-lingu…
AfroDigits: A Community-Driven Spoken Digit Dataset for African Languages
Chris Chinenye Emezue, Sanchit Gandhi, Lewis Tunstall +10
The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address thi…
MasakhaNEWS: News Topic Classification for African languages
David Ifeoluwa Adelani, Marek Masiak, Israel Abebe Azime +62
African languages are severely under-represented in NLP research due to lack of datasets covering several NLP tasks. While there are individual language specific datasets that are…
Adapting to the Low-Resource Double-Bind: Investigating Low-Compute Methods on Low-Resource African Languages
Colin Leong, Herumb Shandilya, Bonaventure F. P. Dossou +10
Many natural language processing (NLP) tasks make use of massively pre-trained language models, which are computationally expensive. However, access to high computational resources…