Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
QPILOTS: Efficient Test-Time Q-Steering for Flow Policies
Yifan Ruan, Chenyang Cao, Andreas Burger +7
Flow-matching and diffusion policies are expressive action generators, but optimizing them with temporal-difference reinforcement learning (RL) remains difficult. Effective policy…
cs.LG2025
Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics
Chenyang Cao, Miguel Rogel-García, Mohamed Nabail +2
Preference-based Reinforcement Learning (PbRL) provides a promising alternative to heuristic reward design in complex robotic environments. However, PbRL often suffers from poor sa…