Showing stat.MLShow all
2 papers · 1 filter
stat.ML2025
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
Gen Li, Yuling Yan
Reinforcement learning with human feedback (RLHF), which learns a reward model from human preference data and then optimizes a policy to favor preferred responses, has emerged as a…
stat.ML2025
Adaptivity and Convergence of Probability Flow ODEs in Diffusion Generative Models
Jiaqi Tang, Yuling Yan
Score-based generative models, which transform noise into data by learning to reverse a diffusion process, have become a cornerstone of modern generative AI. This paper contributes…