3 papers
cs.CL2026
Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition
Kesego Mokgosi, Vukosi Marivate, Sitwala Mundia +3
Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and…
cs.CL2025
Mafoko: Structuring and Building Open Multilingual Terminologies for South African NLP
Vukosi Marivate, Isheanesu Dzingirai, Fiskani Banda +9
The critical lack of structured terminological data for South Africa's official languages hampers progress in multilingual NLP, despite the existence of numerous government and aca…
cs.CL2024
From N-grams to Pre-trained Multilingual Models For Language Identification
Thapelo Sindane, Vukosi Marivate
In this paper, we investigate the use of N-gram models and Large Pre-trained Multilingual models for Language Identification (LID) across 11 South African languages. For N-gram mod…