7 citations · 8 across the 7 of their papers we have counts for
1 paper · 1 filter
Khiem Pham, Quang Nguyen, Tung Nguyen +4
Direct Preference Optimization (DPO) has emerged as a predominant alignment method for diffusion models, facilitating off-policy training without explicit reward modeling. However,…