Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
Bingning Huang, Tu Nguyen, Matthieu Zimmer
Recent advances in reasoning with large language models (LLMs) have shown the effectiveness of Monte Carlo Tree Search (MCTS) for generating high quality intermediate trajectories,…
cs.AI2025
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
Matthieu Zimmer, Xiaotong Ji, Rasul Tutunov +3
Reasoning remains a challenging task for large language models (LLMs), especially within the logically constrained environment of automated theorem proving (ATP), due to sparse rew…