activity
20242026
collaborators

16 papers

cs.CL2026

Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition

Kesego Mokgosi, Vukosi Marivate, Sitwala Mundia +3

Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and…

cs.CL2026

The Annotation Scarcity Paradox in Low-Resource NLP Evaluation: A Decade of Acceleration and Emerging Constraints

Vukosi Marivate

Over the past decade, low-resource natural language processing (NLP) has experienced explosive growth, propelled by cross-lingual transfer, massively multilingual models, and the r…

cs.CL2026

Swivuriso: The South African Next Voices Multilingual Speech Dataset

Vukosi Marivate, Kayode Olaleye, Sitwala Mundia +19

This paper introduces Swivuriso, a 3000-hour multilingual speech dataset developed as part of the African Next Voices project, to support the development and benchmarking of automa…

cs.CL2026

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94

Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…

cs.CL2026

AfriScience-MT: Towards Decolonizing Science in Africa through Text Translation

Idris Abdulmumin, Tajuddeen Gwadabe, Shamsuddeen Hassan Muhammad +11

The dominance of colonial languages in African education and scientific communication limits how hundreds of millions of speakers of African languages access and produce scientific…

cs.CL2026

Temporal Simultaneity Predicts Annotation Quality in Sentiment Corpora

Idris Abdulmumin, Mokgadi Penelope Matloga, Tadesse Destaw Belay +5

Annotation quality is difficult to sustain when campaigns span weeks or months with small annotator pools. We present a Setswana sentiment dataset of 3,565 tweets annotated by thre…