3 papers
stat.ML2025
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
Gen Li, Yuling Yan
Reinforcement learning with human feedback (RLHF), which learns a reward model from human preference data and then optimizes a policy to favor preferred responses, has emerged as a…
cs.LG2025
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions
Gen Li, Yuling Yan
Score-based diffusion models, which generate new data by learning to reverse a diffusion process that perturbs data from the target distribution into noise, have achieved remarkabl…
cs.LG2024
Adapting to Unknown Low-Dimensional Structures in Score-Based Diffusion Models
Gen Li, Yuling Yan
This paper investigates score-based diffusion models when the underlying target distribution is concentrated on or near low-dimensional manifolds within the higher-dimensional spac…