2 papers
cs.AI2026
Game Arena: Strategic LLM Evaluation in Competitive Environments
Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu +59
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena…
cs.AI2026
Measuring Progress Toward AGI: A Cognitive Framework
Ryan Burnell, Yumeya Yamamori, Orhan Firat +10
Despite widespread discussion of AGI, there is no clear framework for measuring progress toward it. This ambiguity fuels subjective claims, makes it difficult to track progress, an…