5 papers · 1 filter
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
Lei Lv, Yunfei Li, Yu Luo +2
Iterative generative policies, such as diffusion models and flow matching, offer superior expressivity for continuous control but complicate Maximum Entropy Reinforcement Learning…
Flow-Based Policy for Online Reinforcement Learning
Lei Lv, Yunfei Li, Yu Luo +4
We present \textbf{FlowRL}, a novel framework for online reinforcement learning that integrates flow-based policy representation with Wasserstein-2-regularized optimization. We arg…
Bidirectional-Reachable Hierarchical Reinforcement Learning with Mutually Responsive Policies
Yu Luo, Fuchun Sun, Tianying Ji +1
Hierarchical reinforcement learning (HRL) addresses complex long-horizon tasks by skillfully decomposing them into subgoals. Therefore, the effectiveness of HRL is greatly influenc…
OMPO: A Unified Framework for RL under Policy and Dynamics Shifts
Yu Luo, Tianying Ji, Fuchun Sun +3
Training reinforcement learning policies using environment interaction data collected from varying policies or dynamics presents a fundamental challenge. Existing works often overl…
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
Yu Luo, Tianying Ji, Fuchun Sun +3
Off-policy reinforcement learning (RL) has achieved notable success in tackling many complex real-world tasks, by leveraging previously collected data for policy learning. However,…