activity
20192026
most citedThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

5 citations · 17 across the 14 of their papers we have counts for

collaborators

14 papers

eess.AS2026

WAXAL: A Large-Scale Multilingual African Language Speech Corpus

Abdoulaye Diack, Perry Nelson, Kwaku Agbesi +40

The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To…

cs.CV2026

Machine Learning for Detection and Severity Estimation of Sweetpotato Weevil Damage in Field and Lab Conditions

Doreen M. Chelangat, Sudi Murindanyi, Bruce Mugizi +6

Sweetpotato weevils (Cylas spp.) are considered among the most destructive pests impacting sweetpotato production, particularly in sub-Saharan Africa. Traditional methods for asses…

cs.CL2025

Benchmarking Automatic Speech Recognition Models for African Languages

Alvin Nahabwe, Sulaiman Kagumire, Denis Musinguzi +5

Automatic speech recognition (ASR) for African languages remains constrained by limited labeled data and the lack of systematic guidance on model selection, data scaling, and decod…

cs.CL2025

Automatic Speech Recognition (ASR) for African Low-Resource Languages: A Systematic Literature Review

Sukairaj Hafiz Imam, Tadesse Destaw Belay, Kedir Yassin Husse +7

ASR has achieved remarkable global progress, yet African low-resource languages remain rigorously underrepresented, producing barriers to digital inclusion across the continent wit…

cs.CL2025

mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks

Luel Hagos Beyene, Vivek Verma, Min Ma +4

Large Language models (LLMs) have demonstrated impressive performance on a wide range of tasks, including in multimodal settings such as speech. However, their evaluation is often…

cs.HC2025

Amplify Initiative: Building A Localized Data Platform for Globalized AI

Qazi Mamunur Rashid, Erin van Liemt, Tiffany Shih +19

Current AI models often fail to account for local context and language, given the predominance of English and Western internet content in their training data. This hinders the glob…