1 paper · 1 filter
Renye Yan, Jikang Cheng, Shikun Sun +7
Despite strong image-generation performance, diffusion models' reconstruction objectives limit alignment with human preferences. RL enables such alignment through explicit rewards.…