1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Khiem Pham, Quang Nguyen, Tung Nguyen +4
Direct Preference Optimization (DPO) has emerged as a predominant alignment method for diffusion models, facilitating off-policy training without explicit reward modeling. However,…