9 papers · 1 filter
GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios
Ke Hu, Shutong Ding, Panxin Tao +2
Generative policies provide expressive and multimodal action distributions, making them attractive for reinforcement learning (RL) in complex continuous-control tasks. Among them,…
Path-Space Mirror Descent for On-Policy Reinforcement Learning under the Generalized Schrödinger Bridge
Yuehu Gong, Zeyuan Wang, Yulin Chen +3
Classical on-policy algorithms such as PPO and mirror descent policy optimization provide stable proximal policy updates through tractable action likelihoods, but are typically ins…
Distributional Reinforcement Learning with Diffusion Bridge Critics
Shutong Ding, Yimiao Zhou, Ke Hu +5
Recent advances in diffusion-based reinforcement learning (RL) methods have demonstrated promising results in a wide range of continuous control tasks. However, existing works in t…
FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning
Shan Zhong, Shutong Ding, He Diao +3
Reliable value estimation serves as the cornerstone of reinforcement learning (RL) by evaluating long-term returns and guiding policy improvement, significantly influencing the con…
GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning
Shutong Ding, Ke Hu, Shan Zhong +5
Recent advances in reinforcement learning (RL) have demonstrated the powerful exploration capabilities and multimodality of generative diffusion-based policies. While substantial p…
Diffusion-based learning framework for Constrained Nonconvex Optimization with Weighted Bootstrapped Refinement
Shutong Ding, Yimiao Zhou, Ke Hu +4
Recent advances in diffusion models show promising potential to accelerate nonconvex problem solving by leveraging their multimodality. However, most existing diffusion-based optim…