2 papers
cs.CV2026
DeDPO: Debiased Direct Preference Optimization for Diffusion Models
Khiem Pham, Quang Nguyen, Tung Nguyen +4
Direct Preference Optimization (DPO) has emerged as a predominant alignment method for diffusion models, facilitating off-policy training without explicit reward modeling. However,…
cs.LG2025
Improving realistic semi-supervised learning with doubly robust estimation
Khiem Pham, Charles Herrmann, Ramin Zabih
A major challenge in Semi-Supervised Learning (SSL) is the limited information available about the class distribution in the unlabeled data. In many real-world applications this ar…