5 papers
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
Xiaotong Ji, Rasul Tutunov, Matthieu Zimmer +1
Reinforcement learning (RL) post-training is a dominant approach for improving the reasoning performance of large language models (LLMs), yet growing evidence suggests that its gai…
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
Matthieu Zimmer, Xiaotong Ji, Tu Nguyen +1
We introduce a novel approach to large language model (LLM) distillation by formulating it as a constrained reinforcement learning problem. While recent work has begun exploring th…
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
Bingning Huang, Tu Nguyen, Matthieu Zimmer
Recent advances in reasoning with large language models (LLMs) have shown the effectiveness of Monte Carlo Tree Search (MCTS) for generating high quality intermediate trajectories,…
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
Matthieu Zimmer, Xiaotong Ji, Rasul Tutunov +3
Reasoning remains a challenging task for large language models (LLMs), especially within the logically constrained environment of automated theorem proving (ATP), due to sparse rew…
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
Xiaotong Ji, Shyam Sundhar Ramesh, Matthieu Zimmer +3
We introduce a novel inference-time alignment approach for LLMs that aims to generate safe responses almost surely, i.e., with probability approaching one. Our approach models the…