papers

Publications (14)

cs.CL2023

ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text Games

Ruoyao Wang, Graham Todd, Eric Yuan +3

In this work, we investigate the capacity of language models to generate explicit, interpretable, and interactive world models of scientific and common-sense reasoning tasks. We op…

q-bio.NC2026

People use fast and flat simulation to reason about new games

Katherine M. Collins, Cedegao E. Zhang, Lionel Wong +6

Games have long been a microcosm for studying planning and reasoning in both natural and artificial intelligence (AI), often focusing on expert-level or even super-human play. But…

cs.AI2025

PuzzleJAX: A Benchmark for Reasoning and Learning

Sam Earle, Graham Todd, Yuchen Li +5

We introduce PuzzleJAX, a GPU-accelerated puzzle game engine and description language designed to support rapid benchmarking of tree search, reinforcement learning, and LLM reasoni…

cs.AI2024

Goals as Reward-Producing Programs

Guy Davidson, Graham Todd, Julian Togelius +2

People are remarkably capable of generating their own goals, beginning with child's play and continuing into adulthood. Despite considerable empirical and computational work on goa…

cs.MA2020

Learning Compositional Negation in Populations of Roth-Erev and Neural Agents

Graham Todd, Shane Steinert-Threlkeld, Christopher Potts

Agent-based models and signalling games are useful tools with which to study the emergence of linguistic communication in a tractable setting. These techniques have been used to st…

cs.CL2026

Evaluating Language Models' Evaluations of Games

Katherine M. Collins, Cedegao E. Zhang, Graham Todd +9

Reasoning is not just about solving problems -- it is also about evaluating which problems are worth solving at all. Evaluations of artificial intelligence (AI) systems primarily f…