1 paper · 1 filter
Yash Savani, Branislav Kveton, Yuchen Liu +5
Flow-GRPO successfully applies reinforcement learning to flow models, but uses uniform credit assignment across all steps. This ignores the temporal structure of diffusion generati…