1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Bingda Tang, Yuhui Zhang, Xiaohan Wang +3
Aligning denoising generative models with human preferences or verifiable rewards remains a key challenge. While policy-gradient online reinforcement learning (RL) offers a princip…