activity
20232026
most citedTARJAMAT: Evaluation of Bard and ChatGPT on Machine Translation of Ten Arabic Varieties

2 citations · 3 across the 13 of their papers we have counts for

collaborators
Showing 2025 · cs.CLShow all

5 papers · 2 filters

cs.CL2025

Are LLMs Good Text Diacritizers? An Arabic and Yoruba Case Study

Hawau Olamide Toyin, Samar Mohamed Magdy, Hanan Aldarmaki

We investigate the effectiveness of large language models (LLMs) for text diacritization in two typologically distinct languages: Arabic and Yoruba. To enable a rigorous evaluation…

cs.CL2025

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis

Hawau Olamide Toyin, Rufael Marew, Humaid Alblooshi +2

We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for…

cs.CL2025

Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset

Fakhraddin Alwajih, Samar M. Magdy, Abdellah El Mekki +34

Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. To address this gap, we introduce PEARL, a…

cs.CL2025

Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs

Fakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy +41

As large language models (LLMs) become increasingly integrated into daily life, ensuring their cultural sensitivity and inclusivity is paramount. We introduce our dataset, a year-l…

cs.CL2025

Jawaher: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking

Samar M. Magdy, Sang Yun Kwon, Fakhraddin Alwajih +3

Recent advancements in instruction fine-tuning, alignment methods such as reinforcement learning from human feedback (RLHF), and optimization techniques like direct preference opti…