3 papers
cs.LG2026
Model-Based Reinforcement Learning with Double Oracle Efficiency in Policy Optimization and Offline Estimation
Haichen Hu, Jian Qian, David Simchi-Levi
Reinforcement learning (RL) in large environments often suffers from severe computational bottlenecks, as conventional regret minimization algorithms require repeated, costly calls…
cs.LG2026
Predicting and improving test-time scaling laws via reward tail-guided search
Muheng Li, Jian Qian, Wenlong Mou
Test-time scaling has emerged as a critical avenue for enhancing the reasoning capabilities of Large Language Models (LLMs). Though the straight-forward ''best-of-'' (BoN) strat…
cs.LG2024
To bootstrap or to rollout? An optimal and adaptive interpolation
Wenlong Mou, Jian Qian
Bootstrapping and rollout are two fundamental principles for value function estimation in reinforcement learning (RL). We introduce a novel class of Bellman operators, called subgr…