1 paper
Yonghoon Dong, Kyungmin Lee, Changyeon Kim +2
Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step sampling process. Recently, Q-l…