From the 1 of 24 linked papers with an AI index.
24 papers
When Do VLMs Help Arabic Manuscript OCR? A Cross-Dataset Study
Moshiur Farazi, Firoj Alam, Abderrahmane Maaradji +3
Vision-language models (VLMs) are increasingly being used for document understanding, yet their role in Arabic and Islamic manuscript recognition remains underexplored. To address…
ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification
Wajdi Zaghouani, Md. Rafiul Biswas, Kholoud Khalil Aldous +1
We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, Egyptian, Algerian, Palestini…
AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Abul Hasnat +3
The paper presents AHA-Memes, a large-scale Arabic hateful meme dataset with fine-grained, multi‑label annotations, and provides baseline evaluations of text, image, and multimodal…
StanceNakba Shared Task: Actor and Topic-Aware Stance Detection in Public Discourse
Kholoud K. Aldous, Md Rafiul Biswas, Mabrouka Bessghaier +3
We present StanceNakba 2026, a shared task on stance detection in polarized social media discourse related to the Palestinian-Israeli conflict, organized as part of Nakba-NLP 2026…
The Generator-Eraser Paradox: Community Guidelines for Responsible LLM-Assisted Dialect Resource Creation
Wajdi Zaghouani
Dialect resources occupy a unique position at the intersection of scientific description, cultural preservation, and computational infrastructure. Large language models offer power…
Toward Responsible and Epistemically Grounded Multilingual LLMs for Computational Social Science and Humanities
Wajdi Zaghouani
Large language models have rapidly evolved in multilingual competence and reasoning capacity, enabling their integration into Social Sciences and Humanities research workflows. Yet…