2 citations · 3 across the 12 of their papers we have counts for
1 paper · 1 filter
Kai Yang, Jian Tao, Jiafei Lyu +6
Using reinforcement learning with human feedback (RLHF) has shown significant promise in fine-tuning diffusion models. Previous methods start by training a reward model that aligns…