2 papers
cs.AI2026
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty +3
We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single ep…
cs.AI2026
Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners
Botos Csaba, Sreejan Kumar, Austin Tudor David Andrews +6
Humans rapidly learn abstract knowledge when encountering novel environments and flexibly deploy this knowledge to guide efficient and intelligent action. Can modern AI systems lea…