5 papers
What should post-training optimize? A test-time scaling law perspective
Muheng Li, Jian Qian, Wenlong Mou
Large language models are increasingly deployed with test-time strategies: sample responses, score them with a reward model or verifier, and return the best. This deployment ru…
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence
Xinquan Chen, Zhenyun Yin, Shan He +38
As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool use, and real environment intera…
Model-Based Reinforcement Learning with Double Oracle Efficiency in Policy Optimization and Offline Estimation
Haichen Hu, Jian Qian, David Simchi-Levi
Reinforcement learning (RL) in large environments often suffers from severe computational bottlenecks, as conventional regret minimization algorithms require repeated, costly calls…
Predicting and improving test-time scaling laws via reward tail-guided search
Muheng Li, Jian Qian, Wenlong Mou
Test-time scaling has emerged as a critical avenue for enhancing the reasoning capabilities of Large Language Models (LLMs). Though the straight-forward ''best-of-'' (BoN) strat…
To bootstrap or to rollout? An optimal and adaptive interpolation
Wenlong Mou, Jian Qian
Bootstrapping and rollout are two fundamental principles for value function estimation in reinforcement learning (RL). We introduce a novel class of Bellman operators, called subgr…