collaborators

6 papers

cs.CL2026

Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation

Amr Keleg, Ahmed Amine Ben Abdallah, Taha Yassine +3

Arabizi refers to Arabic written in Latin script. Although previous studies have shown that the prevalence and usage of Arabizi vary by factors such as region and age group, most N…

cs.CL2026

Constrained CTC Decoding for Efficient Diacritic Restoration

Rufael Marew, Amr Keleg, Hanan Aldarmaki

In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling fine-grained phonological distinc…

cs.CL2026

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94

Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…

cs.CL2026

Curriculum Learning and Pseudo-Labeling Improve the Generalization of Multi-Label Arabic Dialect Identification Models

Ali Mekky, Mohamed El Zeftawy, Lara Hassan +2

Being modeled as a single-label classification task for a long time, recent work has argued that Arabic Dialect Identification (ADI) should be framed as a multi-label classificatio…

cs.CL2025

Revisiting Common Assumptions about Arabic Dialects in NLP

Amr Keleg, Sharon Goldwater, Walid Magdy

Arabic has diverse dialects, where one dialect can be substantially different from the others. In the NLP literature, some assumptions about these dialects are widely adopted (e.g.…

cs.CL2025

LLM Alignment for the Arabs: A Homogenous Culture or Diverse Ones?

Amr Keleg

Large language models (LLMs) have the potential of being useful tools that can automate tasks and assist humans. However, these models are more fluent in English and more aligned w…