1 citations · 2 across the 11 of their papers we have counts for
1 paper · 1 filter
Bingda Tang, Yuhui Zhang, Xiaohan Wang +3
Aligning denoising generative models with human preferences or verifiable rewards remains a key challenge. While policy-gradient online reinforcement learning (RL) offers a princip…