13 papers
Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages
Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida +16
Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpreta…
Aligning Stuttered-Speech Research with End-User Needs: Scoping Review, Survey, and Guidelines
Hawau Olamide Toyin, Mutiah Apampa, Toluwani Aremu +6
Atypical speech is receiving greater attention in speech technology research, but much of this work unfolds with limited interdisciplinary dialogue. For stuttered speech in particu…
Are LLMs Good Text Diacritizers? An Arabic and Yoruba Case Study
Hawau Olamide Toyin, Samar Mohamed Magdy, Hanan Aldarmaki
We investigate the effectiveness of large language models (LLMs) for text diacritization in two typologically distinct languages: Arabic and Yoruba. To enable a rigorous evaluation…
NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
Bashar Talafha, Hawau Olamide Toyin, Peter Sullivan +9
We present the findings of the sixth Nuanced Arabic Dialect Identification (NADI 2025) Shared Task, which focused on Arabic speech dialect processing across three subtasks: spoken…
Voice of a Continent: Mapping Africa's Speech Technology Frontier
AbdelRahim Elmadany, Sang Yun Kwon, Hawau Olamide Toyin +3
Africa's rich linguistic diversity remains significantly underrepresented in speech technologies, creating barriers to digital inclusion. To alleviate this challenge, we systematic…
Towards a Unified Benchmark for Arabic Pronunciation Assessment: Quranic Recitation as Case Study
Yassine El Kheir, Omnia Ibrahim, Amit Meghanani +12
We present a unified benchmark for mispronunciation detection in Modern Standard Arabic (MSA) using Qur'anic recitation as a case study. Our approach lays the groundwork for advanc…