2 papers
cs.LG2026
ARROW: Augmented Replay for RObust World models
Abdulaziz Alyahya, Abdallah Al Siyabi, Markus R. Ernst +3
Continual reinforcement learning challenges agents to acquire new skills while retaining previously learned ones with the goal of improving performance in both past and future task…
cs.AI2025
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
Haidar Khan, Hisham A. Alyahya, Yazeed Alnumay +2
Evaluating the capabilities of Large Language Models (LLMs) has traditionally relied on static benchmark datasets, human assessments, or model-based evaluations - methods that ofte…