5 papers
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
Gen Li, Yuling Yan
Reinforcement learning with human feedback (RLHF), which learns a reward model from human preference data and then optimizes a policy to favor preferred responses, has emerged as a…
Gaussian mixture layers for neural networks
Sinho Chewi, Philippe Rigollet, Yuling Yan
The mean-field theory for two-layer neural networks considers infinitely wide networks that are linearly parameterized by a probability measure over the parameter space. This nonpa…
Adaptivity and Convergence of Probability Flow ODEs in Diffusion Generative Models
Jiaqi Tang, Yuling Yan
Score-based generative models, which transform noise into data by learning to reverse a diffusion process, have become a cornerstone of modern generative AI. This paper contributes…
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions
Gen Li, Yuling Yan
Score-based diffusion models, which generate new data by learning to reverse a diffusion process that perturbs data from the target distribution into noise, have achieved remarkabl…
Adapting to Unknown Low-Dimensional Structures in Score-Based Diffusion Models
Gen Li, Yuling Yan
This paper investigates score-based diffusion models when the underlying target distribution is concentrated on or near low-dimensional manifolds within the higher-dimensional spac…