1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CL2023
ALDi: Quantifying the Arabic Level of Dialectness of Text
Amr Keleg, Sharon Goldwater, Walid Magdy
Transcribed speech and user-generated text in Arabic typically contain a mixture of Modern Standard Arabic (MSA), the standardized language taught in schools, and Dialectal Arabic…
cs.CL2023
Arabic Dialect Identification under Scrutiny: Limitations of Single-label Classification
Amr Keleg, Walid Magdy
Automatic Arabic Dialect Identification (ADI) of text has gained great popularity since it was introduced in the early 2010s. Multiple datasets were developed, and yearly shared ta…
cs.CL2023★ 1 cited
DLAMA: A Framework for Curating Culturally Diverse Facts for Probing the Knowledge of Pretrained Language Models
Amr Keleg, Walid Magdy
A few benchmarking datasets have been released to evaluate the factual knowledge of pretrained language models. These benchmarks (e.g., LAMA, and ParaRel) are mainly developed in E…