8 papers
Bulbul: A Dataset for Dialectal Arabic Speech Recognition
Ahmed Ashraf, Aisha Alansari, Fadel Al Abbas +30
Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets oft…
MoCA-Agent: A Market-of-Claims Code Agent for Financial and Numerical Reasoning
Abdelrahman Abdallah, AbdelRahim A. Elmadany, Sameh Al Natour +3
Financial and tabular question answering requires more than fluent reasoning: answers must be grounded in the exact facts, formulas, units, signs, and scales that support them. A s…
AfroScope: A Framework for Studying the Linguistic Landscape of Africa
Sang Yun Kwon, AbdelRahim Elmadany, Muhammad Abdul-Mageed
Language Identification (LID), the task of determining the language of a given text, is a fundamental preprocessing step that shapes the reliability of downstream NLP applications.…
Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs
Abdellah El Mekki, Samar M. Magdy, Houdaifa Atou +44
Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than Modern Standard Arabic (MSA). Despite this, machine translation (MT) sy…
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology
Peter Sullivan, AbdelRahim Elmadany, Alcides Alcoba Inciarte +1
Dialectal Arabic (DA) speech data vary widely in domain coverage, dialect labeling practices, and recording conditions, complicating cross-dataset comparison and model evaluation.…
NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
Bashar Talafha, Hawau Olamide Toyin, Peter Sullivan +9
We present the findings of the sixth Nuanced Arabic Dialect Identification (NADI 2025) Shared Task, which focused on Arabic speech dialect processing across three subtasks: spoken…