16 papers
Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition
Kesego Mokgosi, Vukosi Marivate, Sitwala Mundia +3
Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and…
The Annotation Scarcity Paradox in Low-Resource NLP Evaluation: A Decade of Acceleration and Emerging Constraints
Vukosi Marivate
Over the past decade, low-resource natural language processing (NLP) has experienced explosive growth, propelled by cross-lingual transfer, massively multilingual models, and the r…
Swivuriso: The South African Next Voices Multilingual Speech Dataset
Vukosi Marivate, Kayode Olaleye, Sitwala Mundia +19
This paper introduces Swivuriso, a 3000-hour multilingual speech dataset developed as part of the African Next Voices project, to support the development and benchmarking of automa…
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
AfriScience-MT: Towards Decolonizing Science in Africa through Text Translation
Idris Abdulmumin, Tajuddeen Gwadabe, Shamsuddeen Hassan Muhammad +11
The dominance of colonial languages in African education and scientific communication limits how hundreds of millions of speakers of African languages access and produce scientific…
Temporal Simultaneity Predicts Annotation Quality in Sentiment Corpora
Idris Abdulmumin, Mokgadi Penelope Matloga, Tadesse Destaw Belay +5
Annotation quality is difficult to sustain when campaigns span weeks or months with small annotator pools. We present a Setswana sentiment dataset of 3,565 tweets annotated by thre…