2 citations · 4 across the 6 of their papers we have counts for
6 papers · 1 filter
NADI 2024: The Fifth Nuanced Arabic Dialect Identification Shared Task
Muhammad Abdul-Mageed, Amr Keleg, AbdelRahim Elmadany +5
We describe the findings of the fifth Nuanced Arabic Dialect Identification Shared Task (NADI 2024). NADI's objective is to help advance SoTA Arabic NLP by providing guidance, data…
Estimating the Level of Dialectness Predicts Interannotator Agreement in Multi-dialect Arabic Datasets
Amr Keleg, Walid Magdy, Sharon Goldwater
On annotating multi-dialect Arabic datasets, it is common to randomly assign the samples across a pool of native Arabic speakers. Recent analyses recommended routing dialectal samp…
ALDi: Quantifying the Arabic Level of Dialectness of Text
Amr Keleg, Sharon Goldwater, Walid Magdy
Transcribed speech and user-generated text in Arabic typically contain a mixture of Modern Standard Arabic (MSA), the standardized language taught in schools, and Dialectal Arabic…
Arabic Dialect Identification under Scrutiny: Limitations of Single-label Classification
Amr Keleg, Walid Magdy
Automatic Arabic Dialect Identification (ADI) of text has gained great popularity since it was introduced in the early 2010s. Multiple datasets were developed, and yearly shared ta…
DLAMA: A Framework for Curating Culturally Diverse Facts for Probing the Knowledge of Pretrained Language Models
Amr Keleg, Walid Magdy
A few benchmarking datasets have been released to evaluate the factual knowledge of pretrained language models. These benchmarks (e.g., LAMA, and ParaRel) are mainly developed in E…
Don't Take it Personally: Analyzing Gender and Age Differences in Ratings of Online Humor
J. A. Meaney, Steven R. Wilson, Luis Chiruzzo +1
Computational humor detection systems rarely model the subjectivity of humor responses, or consider alternative reactions to humor - namely offense. We analyzed a large dataset of…