collaborators

18 papers

cs.CL2026

Constrained CTC Decoding for Efficient Diacritic Restoration

Rufael Marew, Amr Keleg, Hanan Aldarmaki

In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling fine-grained phonological distinc…

cs.CL2026

What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR

Hawau Olamide Toyin, Srinivasan Umesh, Hanan Aldarmaki

ASR systems have been often reported to underperform on atypical speech. An often conflated compounding factor is the existence of two valid transcription references: verbatim (act…

cs.CL2026

Linear Semantic Segmentation for Low-Resource Spoken Dialects

Kirill Chirkunov, Younes Samih, Abed Alhakim Freihat +1

Semantic segmentation is a core component of discourse analysis, yet existing models are primarily developed and evaluated on high-resource written text, limiting their effectivene…

cs.CL2026

Unrequited Emotions: Investigating the Gaps in Motivation and Practice in Speech Emotion Recognition Research

Taryn Wong, Zeerak Talat, Hanan Aldarmaki +1

Critical analyses of emotion recognition technology have raised ethical concerns around task validity and potential downstream impacts, urging researchers to ensure alignment betwe…

cs.CL2026

Aligning Stuttered-Speech Research with End-User Needs: Scoping Review, Survey, and Guidelines

Hawau Olamide Toyin, Mutiah Apampa, Toluwani Aremu +6

Atypical speech is receiving greater attention in speech technology research, but much of this work unfolds with limited interdisciplinary dialogue. For stuttered speech in particu…

cs.CL2026

Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs

Yara Alakeel, Chatrine Qwaider, Hanan Aldarmaki +1

This work investigates how effectively large language models (LLMs) and their tokenization schemes represent and generate Arabic root-pattern morphology, probing whether they captu…