4 papers
Are Arabic Benchmarks Reliable? QIMMA's Quality-First Approach to LLM Evaluation
Leen AlQadi, Ahmed Alzubaidi, Mohammed Alyafeai +6
We present QIMMA, a quality-assured Arabic LLM leaderboard that places systematic benchmark validation at its core. Rather than aggregating existing resources as-is, QIMMA applies…
Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps
Ahmed Alzubaidi, Shaikha Alsuwaidi, Basma El Amel Boussaha +5
This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and spec…
3LM: Bridging Arabic, STEM, and Code through Benchmarking
Basma El Amel Boussaha, Leen AlQadi, Mugariya Farooq +5
Arabic is one of the most widely spoken languages in the world, yet efforts to develop and evaluate Large Language Models (LLMs) for Arabic remain relatively limited. Most existing…
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models
Mouadh Yagoubi, Yasser Dahou, Billel Mokeddem +12
Existing benchmarks have proven effective for assessing the performance of fully trained large language models. However, we find striking differences in the early training stages o…