3 papers
cs.LG2025
Avoiding scaling in RLHF through Preference-based Exploration
Mingyu Chen, Yiding Chen, Wen Sun +1
Reinforcement Learning from Human Feedback (RLHF) has emerged as a pivotal technique for large language model (LLM) alignment. This paper studies the setting of online RLHF and foc…
cs.LG2025
Efficient Controllable Diffusion via Optimal Classifier Guidance
Owen Oertell, Shikun Sun, Yiding Chen +3
The controllable generation of diffusion models aims to steer the model to generate samples that optimize some given objective functions. It is desirable for a variety of applicati…
cs.LG2025
Convergence Of Consistency Model With Multistep Sampling Under General Data Assumptions
Yiding Chen, Yiyi Zhang, Owen Oertell +1
Diffusion models accomplish remarkable success in data generation tasks across various domains. However, the iterative sampling process is computationally expensive. Consistency mo…