3 papers
stat.ML2026
FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models
Yansen Han, Shengyi Liao, Peng Sun +4
Preference alignment for flow and diffusion models now spans online reinforcement learning and offline preference optimization, but the relation between these methods remains uncle…
cs.LG2026
When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood?
Yansen Han, Hongxin Sun, Tao Lin
Flow matching enables likelihood-free training, yet alignment methods increasingly reuse conditional flow matching (CFM) losses as endpoint negative log-likelihoods (NLLs) and thei…
cs.AI2026
Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking
Yansen Han, Shengyi Liao, Yuanxing Zhang +2
Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven upda…