activity
20242026
collaborators

11 papers

cs.CE2026

Hierarchical Constrained Reinforcement Learning with Dynamic Boundary for Spatio-Temporal Vehicle-to-Grid Scheduling

Haoyu Yan, Shutong Ding, Jiebao Zhang +5

The rapid proliferation of Electric Vehicles (EVs) introduces significant spatio-temporal uncertainties into power grids, while Vehicle-to-Grid (V2G) technology offers critical fle…

cs.RO2026

Steering Generative Reinforcement Learning into Stable Robotic Controller

Yixuan Wang, Shutong Ding, Ke Hu +3

Diffusion and flow-based generative policies provide a powerful policy class for reinforcement learning by inducing rich stochastic exploration through iterative action generation.…

cs.LG2026

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios

Ke Hu, Shutong Ding, Panxin Tao +2

Generative policies provide expressive and multimodal action distributions, making them attractive for reinforcement learning (RL) in complex continuous-control tasks. Among them,…

cs.RO2026

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance

Shutong Ding, Zejia Zhong, Zhongyi Wang +4

Recent advances in reinforcement learning (RL) have achieved great successes by leveraging the multimodality and exploration capability of diffusion policies. Among these approache…

cs.LG2026

Path-Space Mirror Descent for On-Policy Reinforcement Learning under the Generalized Schrödinger Bridge

Yuehu Gong, Zeyuan Wang, Yulin Chen +3

Classical on-policy algorithms such as PPO and mirror descent policy optimization provide stable proximal policy updates through tractable action likelihoods, but are typically ins…

cs.LG2026

Diffusion-based learning framework for Constrained Nonconvex Optimization with Weighted Bootstrapped Refinement

Shutong Ding, Yimiao Zhou, Ke Hu +4

Recent advances in diffusion models show promising potential to accelerate nonconvex problem solving by leveraging their multimodality. However, most existing diffusion-based optim…