3 papers
cs.CL2026
Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition
Kesego Mokgosi, Vukosi Marivate, Sitwala Mundia +3
Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and…
cs.LG2026
Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models
Bruce A. Bassett, Amy Rouillard, Sitwala Mundia +8
Background: Large language models (LLMs) are increasingly proposed for diagnostic support, but few evaluations use real-world multimodal inpatient data, particularly in low and mid…
cs.CV2026
From Handwriting to Structured Data: Benchmarking AI Digitisation of Handwritten Forms
Nicholas Pather, Joshua Fouché, Sitwala Mundia +5
Manual digitisation of structured handwritten documents is slow and costly. We benchmark 17 leading frontier multi-modal large language models and open-source models against a very…