10 citations · 18 across the 7 of their papers we have counts for
6 papers · 1 filter
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…
MultiProSE: A Multi-label Arabic Dataset for Propaganda, Sentiment, and Emotion Detection
Lubna Al-Henaki, Hend Al-Khalifa, Abdulmalik Al-Salman +4
Propaganda is a form of persuasion that has been used throughout history with the intention goal of influencing people's opinions through rhetorical and psychological persuasion te…
A Survey of Large Language Models for Arabic Language and its Dialects
Malak Mashaabi, Shahad Al-Khalifa, Hend Al-Khalifa
This survey offers a comprehensive overview of Large Language Models (LLMs) designed for Arabic language and its dialects. It covers key architectures, including encoder-only, deco…
CLEANANERCorp: Identifying and Correcting Incorrect Labels in the ANERcorp Dataset
Mashael Al-Duwais, Hend Al-Khalifa, Abdulmalik Al-Salman
Label errors are a common issue in machine learning datasets, particularly for tasks such as Named Entity Recognition. Such label errors might hurt model training, affect evaluatio…
The Saudi Privacy Policy Dataset
Hend Al-Khalifa, Malak Mashaabi, Ghadi Al-Yahya +1
This paper introduces the Saudi Privacy Policy Dataset, a diverse compilation of Arabic privacy policies from various sectors in Saudi Arabia, annotated according to the 10 princip…
A Panoramic Survey of Natural Language Processing in the Arab World
Kareem Darwish, Nizar Habash, Mourad Abbas +9
The term natural language refers to any system of symbolic communication (spoken, signed or written) without intentional human planning and design. This distinguishes natural langu…