6 papers
Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation
Amr Keleg, Ahmed Amine Ben Abdallah, Taha Yassine +3
Arabizi refers to Arabic written in Latin script. Although previous studies have shown that the prevalence and usage of Arabizi vary by factors such as region and age group, most N…
Constrained CTC Decoding for Efficient Diacritic Restoration
Rufael Marew, Amr Keleg, Hanan Aldarmaki
In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling fine-grained phonological distinc…
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
Curriculum Learning and Pseudo-Labeling Improve the Generalization of Multi-Label Arabic Dialect Identification Models
Ali Mekky, Mohamed El Zeftawy, Lara Hassan +2
Being modeled as a single-label classification task for a long time, recent work has argued that Arabic Dialect Identification (ADI) should be framed as a multi-label classificatio…
Revisiting Common Assumptions about Arabic Dialects in NLP
Amr Keleg, Sharon Goldwater, Walid Magdy
Arabic has diverse dialects, where one dialect can be substantially different from the others. In the NLP literature, some assumptions about these dialects are widely adopted (e.g.…
LLM Alignment for the Arabs: A Homogenous Culture or Diverse Ones?
Amr Keleg
Large language models (LLMs) have the potential of being useful tools that can automate tasks and assist humans. However, these models are more fluent in English and more aligned w…