1 paper
Chenye Meng, Zejian Li, Zhongni Liu +9
Post-training alignment of diffusion models relies on simplified signals, such as scalar rewards or binary preferences. This limits alignment with complex human expertise, which is…