1 paper
Yifan Ruan, Chenyang Cao, Andreas Burger +7
Flow-matching and diffusion policies are expressive action generators, but optimizing them with temporal-difference reinforcement learning (RL) remains difficult. Effective policy…