2 papers
cs.CL2026
Are Arabic Benchmarks Reliable? QIMMA's Quality-First Approach to LLM Evaluation
Leen AlQadi, Ahmed Alzubaidi, Mohammed Alyafeai +6
We present QIMMA, a quality-assured Arabic LLM leaderboard that places systematic benchmark validation at its core. Rather than aggregating existing resources as-is, QIMMA applies…
cs.CL2025
Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps
Ahmed Alzubaidi, Shaikha Alsuwaidi, Basma El Amel Boussaha +5
This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and spec…