4 citations · 4 across the 2 of their papers we have counts for
4 papers
Divergence Minimization Preference Optimization for Diffusion Model Alignment
Binxu Li, Minkai Xu, Jiaqi Han +2
Diffusion models have achieved remarkable success in generating realistic and versatile images from text prompts. Inspired by the recent advancements of language models, there is a…
Discrete Diffusion Trajectory Alignment via Stepwise Decomposition
Jiaqi Han, Austin Wang, Minkai Xu +6
Discrete diffusion models have demonstrated great promise in modeling various sequence data, ranging from human language to biological sequences. Inspired by the success of RL in l…
Personalized Preference Fine-tuning of Diffusion Models
Meihua Dang, Anikait Singh, Linqi Zhou +2
RLHF techniques like DPO can significantly improve the generation quality of text-to-image diffusion models. However, these methods optimize for a single reward that aligns model g…
Diffusion Model Alignment Using Direct Preference Optimization
Bram Wallace, Meihua Dang, Rafael Rafailov +7
Large language models (LLMs) are fine-tuned using human comparison data with Reinforcement Learning from Human Feedback (RLHF) methods to make them better aligned with users' prefe…