collaborators

7 papers

cs.CL2026

Bulbul: A Dataset for Dialectal Arabic Speech Recognition

Ahmed Ashraf, Aisha Alansari, Fadel Al Abbas +30

Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets oft…

cs.CL2026

Character Iconicity vs. Arbitrariness: An Arabic NLP Perspective

Dorieh Alomari, Irfan Ahmad, Maged S. Al-shaibani

Arabic script uses 28 letters, many of which share a common base shape (rasm) and are distinguished only by dot placement. Because early Arabic manuscripts were written without dot…

cs.CL2025

MeXtract: Light-Weight Metadata Extraction from Scientific Papers

Zaid Alyafeai, Maged S. Al-Shaibani, Bernard Ghanem

Metadata plays a critical role in indexing, documenting, and analyzing scientific literature, yet extracting it accurately and efficiently remains a challenging task. Traditional a…

cs.CL2025

Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset

Fakhraddin Alwajih, Samar M. Magdy, Abdellah El Mekki +34

Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. To address this gap, we introduce PEARL, a…

cs.CL2025

MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMs

Zaid Alyafeai, Maged S. Al-Shaibani, Bernard Ghanem

Metadata extraction is essential for cataloging and preserving datasets, enabling effective research discovery and reproducibility, especially given the current exponential growth…

cs.CL2025

The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text

Maged S. Al-Shaibani, Moataz Ahmed

Large Language Models (LLMs) have achieved unprecedented capabilities in generating human-like text, posing subtle yet significant challenges for information integrity across criti…