16 papers
People use fast and flat simulation to reason about new games
Katherine M. Collins, Cedegao E. Zhang, Lionel Wong +6
Games have long been a microcosm for studying planning and reasoning in both natural and artificial intelligence (AI), often focusing on expert-level or even super-human play. But…
A Matter of Interest: Understanding Interestingness of Math Problems in Humans and Language Models
Shubhra Mishra, Yuka Machino, Gabriel Poesia +9
The evolution of mathematics is shaped importantly by interestingness: researchers choose which problems to pursue, and students choose which problems to engage with, based on expe…
Evaluating Language Models' Evaluations of Games
Katherine M. Collins, Cedegao E. Zhang, Graham Todd +9
Reasoning is not just about solving problems -- it is also about evaluating which problems are worth solving at all. Evaluations of artificial intelligence (AI) systems primarily f…
Medical Model Synthesis Architectures: A Case Study
Katherine M. Collins, Marlene Berke, Ilia Sucholutsky +6
Medicine is rife with high-stakes uncertainty. Doctors routinely make clinical judgments and decisions that juggle many fundamental unknowns, like predictions about what might be c…
ExoPredicator: Learning Abstract Models of Dynamic Worlds for Robot Planning
Yichao Liang, Dat Nguyen, Cambridge Yang +7
Long-horizon embodied planning is challenging because the world does not only change through an agent's actions: exogenous processes (e.g., water heating, dominoes cascading) unfol…
Data for Mathematical Copilots: Better Ways of Presenting Proofs for Machine Learning
Simon Frieder, Jonas Bayer, Sam Looi +13
The datasets and benchmarks commonly used to train and evaluate the mathematical capabilities of AI-based mathematical copilots (primarily large language models) exhibit several sh…