1 paper
Meihua Dang, Anikait Singh, Linqi Zhou +2
RLHF techniques like DPO can significantly improve the generation quality of text-to-image diffusion models. However, these methods optimize for a single reward that aligns model g…