4 papers
Evaluating Language Models' Evaluations of Games
Katherine M. Collins, Cedegao E. Zhang, Graham Todd +9
Reasoning is not just about solving problems -- it is also about evaluating which problems are worth solving at all. Evaluations of artificial intelligence (AI) systems primarily f…
AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
Lance Ying, Ryan Truong, Prafull Sharma +9
Rigorously evaluating machine intelligence against the broad spectrum of human general intelligence has become increasingly important and challenging in this era of rapid technolog…
Digital Red Queen: Adversarial Program Evolution in Core War with LLMs
Akarsh Kumar, Ryan Bahlous-Boldi, Prafull Sharma +4
Large language models (LLMs) are increasingly being used to evolve solutions to problems in many domains, in a process inspired by biological evolution. However, unlike biological…
Assessing Adaptive World Models in Machines with Novel Games
Lance Ying, Katherine M. Collins, Prafull Sharma +11
Human intelligence exhibits a remarkable capacity for rapid adaptation and effective problem-solving in novel and unfamiliar contexts. We argue that this profound adaptability is f…