activity
20202023
most citedThe GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

52 citations · 92 across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

19 papers · 1 filter

cs.CL2023

AfriNames: Most ASR models "butcher" African Names

Tobi Olatunji, Tejumade Afonja, Bonaventure F. P. Dossou +4

Useful conversational agents must accurately capture named entities to minimize error for downstream tasks, for example, asking a voice assistant to play a track from a certain art…

cs.CL20231 cited

MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African Languages

Cheikh M. Bamba Dione, David Adelani, Peter Nabende +41

In this paper, we present MasakhaPOS, the largest part-of-speech (POS) dataset for 20 typologically diverse African languages. We discuss the challenges in annotating POS for these…

cs.CL2023

AfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages

Odunayo Ogundepo, Tajuddeen R. Gwadabe, Clara E. Rivera +49

African languages have far less in-language content available digitally, making it challenging for question answering systems to satisfy the information needs of users. Cross-lingu…

cs.CL20231 cited

AfroDigits: A Community-Driven Spoken Digit Dataset for African Languages

Chris Chinenye Emezue, Sanchit Gandhi, Lewis Tunstall +10

The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address thi…

cs.CL2023

MasakhaNEWS: News Topic Classification for African languages

David Ifeoluwa Adelani, Marek Masiak, Israel Abebe Azime +62

African languages are severely under-represented in NLP research due to lack of datasets covering several NLP tasks. While there are individual language specific datasets that are…

cs.CL20232 cited

Adapting to the Low-Resource Double-Bind: Investigating Low-Compute Methods on Low-Resource African Languages

Colin Leong, Herumb Shandilya, Bonaventure F. P. Dossou +10

Many natural language processing (NLP) tasks make use of massively pre-trained language models, which are computationally expensive. However, access to high computational resources…