3 papers
cs.CL2025
BALSAM: A Platform for Benchmarking Arabic Large Language Models
Rawan Al-Matham, Kareem Darwish, Raghad Al-Rasheed +40
The impressive advancement of Large Language Models (LLMs) in English has not been matched across all languages. In particular, LLM performance in Arabic lags behind, due to data s…
cs.CL2025
A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment
Khalid N. Elmadani, Nizar Habash, Hanada Taha-Thomure
This paper introduces the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 69,441 sen…
cs.CL2025
Guidelines for Fine-grained Sentence-level Arabic Readability Annotation
Nizar Habash, Hanada Taha-Thomure, Khalid N. Elmadani +2
This paper presents the annotation guidelines of the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale resource for fine-grained sentence-level readability asses…