4 papers
Bulbul: A Dataset for Dialectal Arabic Speech Recognition
Ahmed Ashraf, Aisha Alansari, Fadel Al Abbas +30
Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets oft…
IQRA 2026: Interspeech Challenge on Automatic Pronunciation Assessment for Modern Standard Arabic (MSA)
Yassine El Kheir, Amit Meghanani, Mostafa Shahin +5
We present the findings of the second edition of the IQRA Interspeech Challenge, a challenge on automatic Mispronunciation Detection and Diagnosis (MDD) for Modern Standard Arabic…
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset
Fakhraddin Alwajih, Samar M. Magdy, Abdellah El Mekki +34
Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. To address this gap, we introduce PEARL, a…
Towards a Unified Benchmark for Arabic Pronunciation Assessment: Quranic Recitation as Case Study
Yassine El Kheir, Omnia Ibrahim, Amit Meghanani +12
We present a unified benchmark for mispronunciation detection in Modern Standard Arabic (MSA) using Qur'anic recitation as a case study. Our approach lays the groundwork for advanc…