continuous-time rl 1discrete diffusion models 1masked diffusion language models 1policy gradient methods 1trajectory subsampling 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.LG2026
A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models
Zikun Zhang, Jiayuan Sheng, David D. Yao +1
The paper introduces a continuous‑time reinforcement‑learning framework for fine‑tuning discrete diffusion models, especially masked diffusion language models, by modeling state dy…
cs.AI2026
Improved techniques for fine-tuning flow models via adjoint matching: a deterministic control pipeline
Zhengyi Guo, Jiayuan Sheng, David D. Yao +1
We propose a deterministic adjoint matching framework that formulates human preference alignment for flow-based generative models as an optimal control problem over velocity fields…
cs.LG2025
Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
Jiayuan Sheng, Hanyang Zhao, Haoxian Chen +2
Reinforcement Learning from Human Feedback (RLHF) is increasingly used to fine-tune diffusion models, but a key challenge arises from the mismatch between stochastic samplers used…