From the 1 of 22 linked papers with an AI index.
22 papers
ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification
Wajdi Zaghouani, Md. Rafiul Biswas, Kholoud Khalil Aldous +1
We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, Egyptian, Algerian, Palestini…
AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Abul Hasnat +3
The paper presents AHA-Memes, a large-scale Arabic hateful meme dataset with fine-grained, multi‑label annotations, and provides baseline evaluations of text, image, and multimodal…
StanceNakba Shared Task: Actor and Topic-Aware Stance Detection in Public Discourse
Kholoud K. Aldous, Md Rafiul Biswas, Mabrouka Bessghaier +3
We present StanceNakba 2026, a shared task on stance detection in polarized social media discourse related to the Palestinian-Israeli conflict, organized as part of Nakba-NLP 2026…
The Generator-Eraser Paradox: Community Guidelines for Responsible LLM-Assisted Dialect Resource Creation
Wajdi Zaghouani
Dialect resources occupy a unique position at the intersection of scientific description, cultural preservation, and computational infrastructure. Large language models offer power…
Toward Responsible and Epistemically Grounded Multilingual LLMs for Computational Social Science and Humanities
Wajdi Zaghouani
Large language models have rapidly evolved in multilingual competence and reasoning capacity, enabling their integration into Social Sciences and Humanities research workflows. Yet…
Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese
Wajdi Zaghouani, Kholoud K. Aldous, Yicheng Gao
When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break down. These systems struggl…