2 papers
cs.RO2026
FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation
Kexin Shi, Junyao Shi, Poorvi Hebbar +5
Real-world reinforcement learning for robotic manipulation remains challenging, and this difficulty is amplified for flow matching policies: applying policy gradient methods to the…
cs.LG2025
Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency model
Jing Zhang, Linjiajie Fang, Kexin Shi +2
``Distribution shift'' is the main obstacle to the success of offline reinforcement learning. A learning policy may take actions beyond the behavior policy's knowledge, referred to…