6 papers
StanceNakba Shared Task: Actor and Topic-Aware Stance Detection in Public Discourse
Kholoud K. Aldous, Md Rafiul Biswas, Mabrouka Bessghaier +3
We present StanceNakba 2026, a shared task on stance detection in polarized social media discourse related to the Palestinian-Israeli conflict, organized as part of Nakba-NLP 2026…
KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models
Wajdi Zaghouani, Shimaa Amer Ibrahim, Aruzhan Muratbek +2
Kazakh is underrepresented in resources for evaluating the safety behavior of large language models. We present KZ-SafetyPrompts, a Kazakh prompt dataset for safety evaluation acro…
ClimateChat-300K: A Multi-Modal Facebook Dataset for Understanding Diverse Perspectives in Climate Communication
Wajdi Zaghouani, Md. Rafiul Biswas, Mabrouka Bessghaier +2
We present ClimateChat-300K, a large-scale dataset of 299,329 public Facebook posts about climate change collected between May 2020 and May 2024 through the CrowdTangle platform. T…
Audience Engagement with Arabic Women's Social Empowerment and Wellbeing: A Decadal Corpus
Wajdi Zaghouani, Mabrouka Bessghaier, MD. Rafiul Biswas +1
This paper presents the Arabic Women and Society Corpus, a ten year collection of 252,487 public Arabic Facebook posts related to women's empowerment and social wellbeing. The corp…
ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination
Wajdi Zaghouani, Shimaa Amer Ibrahim, Mabrouka Bessghaier +1
We present ArabDiscrim, a decade-long lexical resource and corpus of 293K public Arabic Facebook posts (2014--2024) discussing racism and discrimination. Unlike existing Twitter-ce…
JobArabi: An Arabic Corpus and Analysis of Job Announcements from Social Media
Wajdi Zaghouani, Shimaa Amer Ibrahim, Mabrouka Bessghaier +1
This paper introduces JobArabi, a large-scale corpus of Arabic job announcements collected from social media between January 2024 and October 2025. The dataset contains 20,528 publ…