2 papers
cs.AI2026
Game Arena: Strategic LLM Evaluation in Competitive Environments
Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu +59
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena…
cs.AI2025
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation
D. Sculley, Will Cukierski, Phil Culliton +8
In this position paper, we observe that empirical evaluation in Generative AI is at a crisis point since traditional ML evaluation and benchmarking strategies are insufficient to m…