Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Path-Space Mirror Descent for On-Policy Reinforcement Learning under the Generalized Schrödinger Bridge
Yuehu Gong, Zeyuan Wang, Yulin Chen +3
Classical on-policy algorithms such as PPO and mirror descent policy optimization provide stable proximal policy updates through tractable action likelihoods, but are typically ins…
cs.LG2026
Distributional Reinforcement Learning with Diffusion Bridge Critics
Shutong Ding, Yimiao Zhou, Ke Hu +5
Recent advances in diffusion-based reinforcement learning (RL) methods have demonstrated promising results in a wide range of continuous control tasks. However, existing works in t…
cs.LG2025
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
Zeyuan Wang, Da Li, Yulin Chen +4
We introduce a one-step generative policy for offline reinforcement learning that maps noise directly to actions via a residual reformulation of MeanFlow, making it compatible with…