11 papers
Formal Disco: Scalable Open-Ended Generation of Formally Verified Programs
Gabriel Poesia, Simon Henniger, Tzu-Han Hsu +2
The cost of producing code is rapidly diminishing with increasingly capable AI agents, while quality assurance of generated programs has not kept pace. Formal verification provides…
A Matter of Interest: Understanding Interestingness of Math Problems in Humans and Language Models
Shubhra Mishra, Yuka Machino, Gabriel Poesia +9
The evolution of mathematics is shaped importantly by interestingness: researchers choose which problems to pursue, and students choose which problems to engage with, based on expe…
The Token Games: Evaluating Language Model Reasoning with Puzzle Duels
Simon Henniger, Gabriel Poesia
Evaluating the reasoning capabilities of Large Language Models is increasingly challenging as models improve. Human curation of hard questions is highly expensive, especially in re…
Learning to Rank the Initial Branching Order of SAT Solvers
Arvid Eriksson, Gabriel Poesia, Roman Bresson +2
Finding good branching orders is key to solving SAT problems efficiently, but finding such branching orders is a difficult problem. Using a learning based approach to predict a goo…
Data for Mathematical Copilots: Better Ways of Presenting Proofs for Machine Learning
Simon Frieder, Jonas Bayer, Sam Looi +13
The datasets and benchmarks commonly used to train and evaluate the mathematical capabilities of AI-based mathematical copilots (primarily large language models) exhibit several sh…
Code-enabled language models can outperform reasoning models on diverse tasks
Cedegao E. Zhang, Cédric Colas, Gabriel Poesia +2
Reasoning models (RMs), language models (LMs) trained with reinforcement learning to produce long-form natural language reasoning, have been remarkably successful, but they still r…