6 papers
SR-OPSD: Self-Referenced On-Policy Self-Distillation
Zhuo Sun, Entong Li, Yanlong Zhao +7
On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, providing a useful complement to re…
Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization
Xiaoyuan Cheng, Wenxuan Yuan, Zhancun Mu +5
Model-based reinforcement learning (RL) can be effectively supported at scale through the use of world models. However, in practice, scaling such approaches remains fundamentally l…
Outlier-robust Diffusion Posterior Sampling for Bayesian Inverse Problems
Yiming Yang, Xiaoyuan Cheng, Yi He +3
Diffusion models have emerged as powerful learned priors for Bayesian inverse problems (BIPs). Diffusion-based solvers rely on a presumed likelihood for the observations in BIPs to…
Fisher Decorator: Refining Flow Policy via a Local Transport Map
Xiaoyuan Cheng, Haoyu Wang, Wenxuan Yuan +4
Recent advances in flow-based offline reinforcement learning (RL) have achieved strong performance by parameterizing policies via flow matching. However, they still face critical t…
How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?
Xiaoyuan Cheng, Wenxuan Yuan, Boyang Li +7
Diffusion policy sampling enables reinforcement learning (RL) to represent multimodal action distributions beyond suboptimal unimodal Gaussian policies. However, existing diffusion…
Information Shapes Koopman Representation
Xiaoyuan Cheng, Wenxuan Yuan, Yiming Yang +4
The Koopman operator provides a powerful framework for modeling dynamical systems and has attracted growing interest from the machine learning community. However, its infinite-dime…