1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Ruihang Li, Mengde Xu, Shuyang Gu +4
Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degr…