continuous-time rl 1discrete diffusion models 1masked diffusion language models 1policy gradient methods 1trajectory subsampling 1
From the 1 of 12 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Improved techniques for fine-tuning flow models via adjoint matching: a deterministic control pipeline
Zhengyi Guo, Jiayuan Sheng, David D. Yao +1
We propose a deterministic adjoint matching framework that formulates human preference alignment for flow-based generative models as an optimal control problem over velocity fields…
cs.AI2025
RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization
Hanyang Zhao, Genta Indra Winata, Anirban Das +4
Recently, numerous preference optimization algorithms have been introduced as extensions to the Direct Preference Optimization (DPO) family. While these methods have successfully a…