11 papers
Reinforcement Learning with Action-Triggered Observations
Alexander Ryabchenko, Wenlong Mou
We introduce Action-Triggered Sporadically Traceable Markov Decision Processes (ATST-MDPs), a reinforcement learning framework for partial observability in which full state observa…
What should post-training optimize? A test-time scaling law perspective
Muheng Li, Jian Qian, Wenlong Mou
Large language models are increasingly deployed with test-time strategies: sample responses, score them with a reward model or verifier, and return the best. This deployment ru…
Provable imitation learning for control of instability in partially-observed Vlasov--Poisson equations
Xiaofan Xia, Qin Li, Wenlong Mou
We consider the stabilization of Vlasov--Poisson plasma dynamics, a central control problem in nuclear fusion. Our focus is the gap between what an ideal controller would use and w…
Continuous-time reinforcement learning: ellipticity enables model-free value function approximation
Wenlong Mou
We study off-policy reinforcement learning for controlling continuous-time Markov diffusion processes with discrete-time observations and actions. We consider model-free algorithms…
Revisiting the Constant Stepsize Stochastic Approximation with Decision-Dependent Markovian Noise
Hadi Hadavi, Wenlong Mou, Sergey Samsonov +1
We revisit the convergence analysis of constant stepsize stochastic approximation (SA) with decision-dependent Markovian noise, with a focus on characterizing the stationary bias a…
Predicting and improving test-time scaling laws via reward tail-guided search
Muheng Li, Jian Qian, Wenlong Mou
Test-time scaling has emerged as a critical avenue for enhancing the reasoning capabilities of Large Language Models (LLMs). Though the straight-forward ''best-of-'' (BoN) strat…