collaborators

5 papers

cs.LG2026

Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening

Xiaotong Ji, Rasul Tutunov, Matthieu Zimmer +1

Reinforcement learning (RL) post-training is a dominant approach for improving the reasoning performance of large language models (LLMs), yet growing evidence suggests that its gai…

cs.LG2025

Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective

Matthieu Zimmer, Xiaotong Ji, Tu Nguyen +1

We introduce a novel approach to large language model (LLM) distillation by formulating it as a constrained reinforcement learning problem. While recent work has begun exploring th…

cs.AI2025

Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning

Bingning Huang, Tu Nguyen, Matthieu Zimmer

Recent advances in reasoning with large language models (LLMs) have shown the effectiveness of Monte Carlo Tree Search (MCTS) for generating high quality intermediate trajectories,…

cs.AI2025

Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving

Matthieu Zimmer, Xiaotong Ji, Rasul Tutunov +3

Reasoning remains a challenging task for large language models (LLMs), especially within the logically constrained environment of automated theorem proving (ATP), due to sparse rew…

cs.LG2025

On Almost Surely Safe Alignment of Large Language Models at Inference-Time

Xiaotong Ji, Shyam Sundhar Ramesh, Matthieu Zimmer +3

We introduce a novel inference-time alignment approach for LLMs that aims to generate safe responses almost surely, i.e., with probability approaching one. Our approach models the…