3 citations · 3 across the 3 of their papers we have counts for
6 papers · 1 filter
Swivuriso: The South African Next Voices Multilingual Speech Dataset
Vukosi Marivate, Kayode Olaleye, Sitwala Mundia +19
This paper introduces Swivuriso, a 3000-hour multilingual speech dataset developed as part of the African Next Voices project, to support the development and benchmarking of automa…
Mafoko: Structuring and Building Open Multilingual Terminologies for South African NLP
Vukosi Marivate, Isheanesu Dzingirai, Fiskani Banda +9
The critical lack of structured terminological data for South Africa's official languages hampers progress in multilingual NLP, despite the existence of numerous government and aca…
Prompting Towards Alleviating Code-Switched Data Scarcity in Under-Resourced Languages with GPT as a Pivot
Michelle Terblanche, Kayode Olaleye, Vukosi Marivate
Many multilingual communities, including numerous in Africa, frequently engage in code-switching during conversations. This behaviour stresses the need for natural language process…
YFACC: A Yorùbá speech-image dataset for cross-lingual keyword localisation through visual grounding
Kayode Olaleye, Dan Oneata, Herman Kamper
Visually grounded speech (VGS) models are trained on images paired with unlabelled spoken captions. Such models could be used to build speech systems in settings where it is imposs…
Attention-Based Keyword Localisation in Speech using Visual Grounding
Kayode Olaleye, Herman Kamper
Visually grounded speech models learn from images paired with spoken captions. By tagging images with soft text labels using a trained visual classifier with a fixed vocabulary, pr…
Towards localisation of keywords in speech using weak supervision
Kayode Olaleye, Benjamin van Niekerk, Herman Kamper
Developments in weakly supervised and self-supervised models could enable speech technology in low-resource settings where full transcriptions are not available. We consider whethe…