8 papers
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
Mohammad Amer Khalil, Raghad Nahas, Ahmad Nassar +1
Sign language is the primary approach of communication for the Deaf and Hard-of-Hearing (DHH) community. While there are numerous benchmarks for high-resource sign languages, low-r…
AraModernBERT: Transtokenized Initialization and Long-Context Encoder Modeling for Arabic
Omar Elshehy, Omer Nacar, Abdelbasset Djamai +3
Encoder-only transformer models remain widely used for discriminative NLP tasks, yet recent architectural advances have largely focused on English. In this work, we present AraMode…
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
Abdul Aziz Snoubara, Baraa Al_Maradni, Haya Al_Naal +4
Speech-based AI educational applications have gained significant interest in recent years, particularly for children. However, children speech research remains limited due to the l…
ASCAT: An Arabic Scientific Corpus and Benchmark for Advanced Translation Evaluation
Serry Sibaee, Khloud Al Jallad, Zineb Yousfi +4
We present ASCAT (Arabic Scientific Corpus for Advanced Translation), a high-quality English-Arabic parallel benchmark corpus designed for scientific translation evaluation constru…
Arabic Little STT: Arabic Children Speech Recognition Dataset
Mouhand Alkadri, Dania Desouki, Khloud Al Jallad
The performance of Artificial Intelligence (AI) systems fundamentally depends on high-quality training data. However, low-resource languages like Arabic suffer from severe data sca…
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi
Natural Language Understanding (NLU) is a basic task in Natural Language Processing (NLP). The evaluation of NLU capabilities has become a trending research topic that attracts res…