4 papers · 1 filter
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
Tu Nguyen, Matthieu Zimmer, Rasul Tutunov +2
A recurring pattern in "reasoning without training" is that base LLMs already assign non-trivial probability mass to correct multi-step solutions; the bottleneck is locating these…
Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning
Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer +2
Lean is increasingly used to judge natural-language mathematical answers, but its signal is partial: many answers never formalize, and a failed proof may reflect an ill-typed state…
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
Matthieu Zimmer, Xiaotong Ji, Rasul Tutunov +3
Reasoning remains a challenging task for large language models (LLMs), especially within the logically constrained environment of automated theorem proving (ATP), due to sparse rew…
Robust Probabilistic Model Checking with Continuous Reward Domains
Xiaotong Ji, Hanchun Wang, Antonio Filieri +1
Probabilistic model checking traditionally verifies properties on the expected value of a measure of interest. This restriction may fail to capture the quality of service of a sign…