2 papers
cs.CL2025
Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps
Ahmed Alzubaidi, Shaikha Alsuwaidi, Basma El Amel Boussaha +5
This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and spec…
cs.CL2025
3LM: Bridging Arabic, STEM, and Code through Benchmarking
Basma El Amel Boussaha, Leen AlQadi, Mugariya Farooq +5
Arabic is one of the most widely spoken languages in the world, yet efforts to develop and evaluate Large Language Models (LLMs) for Arabic remain relatively limited. Most existing…