Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Constrained CTC Decoding for Efficient Diacritic Restoration
Rufael Marew, Amr Keleg, Hanan Aldarmaki
In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling fine-grained phonological distinc…
cs.CL2025
Clinical Annotations for Automatic Stuttering Severity Assessment
Ana Rita Valente, Rufael Marew, Hawau Olamide Toyin +6
Stuttering is a complex disorder that requires specialized expertise for effective assessment and treatment. This paper presents an effort to enhance the FluencyBank dataset with a…
cs.CL2025
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
Hawau Olamide Toyin, Rufael Marew, Humaid Alblooshi +2
We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for…