3 papers
cs.LG2026
Path-Space Mirror Descent for On-Policy Reinforcement Learning under the Generalized Schrödinger Bridge
Yuehu Gong, Zeyuan Wang, Yulin Chen +3
Classical on-policy algorithms such as PPO and mirror descent policy optimization provide stable proximal policy updates through tractable action likelihoods, but are typically ins…
cs.LG2026
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
Zeyuan Wang, Da Li, Yulin Chen +6
Online off-policy reinforcement learning (RL) is shaped by two coupled choices: the policy class and the update rule. Gaussian policies are fast and have tractable entropy, but str…
cs.LG2025
Bellman Error Centering
Xingguo Chen, Yu Gong, Shangdong Yang +1
This paper revisits the recently proposed reward centering algorithms including simple reward centering (SRC) and value-based reward centering (VRC), and points out that SRC is ind…