2 papers
cs.CL2025
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
Hisham A. Alyahya, Haidar Khan, Yazeed Alnumay +2
We introduce ZeroSumEval, a dynamic, competition-based, and evolving evaluation framework for Large Language Models (LLMs) that leverages competitive games. ZeroSumEval encompasses…
cs.AI2024
How to Correctly do Semantic Backpropagation on Language-based Agentic Systems
Wenyi Wang, Hisham A. Alyahya, Dylan R. Ashley +4
Language-based agentic systems have shown great promise in recent years, transitioning from solving small-scale research problems to being deployed in challenging real-world tasks.…