3 papers
stat.ML2025
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
Gen Li, Yuling Yan
Reinforcement learning with human feedback (RLHF), which learns a reward model from human preference data and then optimizes a policy to favor preferred responses, has emerged as a…
cs.LG2024
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions
Gen Li, Yuling Yan
Score-based diffusion models, which generate new data by learning to reverse a diffusion process that perturbs data from the target distribution into noise, have achieved remarkabl…
cs.LG2024
A Score-Based Density Formula, with Applications in Diffusion Generative Models
Gen Li, Yuling Yan
Score-based generative models (SGMs) have revolutionized the field of generative modeling, achieving unprecedented success in generating realistic and diverse content. Despite empi…