4 papers
AfriVox-v2: A Domain-Verticalized Benchmark for In-the-Wild African Speech Recognition
Busayo Awobade, Gabrial Zencha Ashungafac, Tobi Olatunji
Recent large language models (LLMs) show strong speech recognition and translation capabilities for high-resource languages. However, African languages remain dramatically underrep…
AfriSpeech-MultiBench: A Verticalized Multidomain Multicountry Benchmark Suite for African Accented English ASR
Gabrial Zencha Ashungafac, Mardhiyah Sanni, Busayo Awobade +2
Recent advances in speech-enabled AI, including Google's NotebookLM and OpenAI's speech-to-speech API, are driving widespread interest in voice interfaces globally. Despite this mo…
From Scarcity to Efficiency: Investigating the Effects of Data Augmentation on African Machine Translation
Mardiyyah Oduwole, Oluwatosin Olajide, Jamiu Suleiman +9
The linguistic diversity across the African continent presents different challenges and opportunities for machine translation. This study explores the effects of data augmentation…
The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages
Chris Emezue, NaijaVoices Community, Busayo Awobade +8
The development of high-performing, robust, and reliable speech technologies depends on large, high-quality datasets. However, African languages -- including our focus, Igbo, Hausa…