2 papers
cs.CL2026
Swivuriso: The South African Next Voices Multilingual Speech Dataset
Vukosi Marivate, Kayode Olaleye, Sitwala Mundia +19
This paper introduces Swivuriso, a 3000-hour multilingual speech dataset developed as part of the African Next Voices project, to support the development and benchmarking of automa…
cs.CL2025
The Esethu Framework: Reimagining Sustainable Dataset Governance and Curation for Low-Resource Languages
Jenalea Rajab, Anuoluwapo Aremu, Everlyn Asiko Chimoto +12
This paper presents the Esethu Framework, a sustainable data curation framework specifically designed to empower local communities and ensure equitable benefit-sharing from their l…