5 papers
Efficient Sampling with Discrete Diffusion Models: Sharp and Adaptive Guarantees
Daniil Dmitriev, Zhihan Huang, Yuting Wei
Diffusion models over discrete spaces have recently shown striking empirical success, yet their theoretical foundations remain incomplete. In this paper, we study the sampling effi…
On the Emergence of Implicit Curriculum in RLVR Learning Dynamics
Yu Huang, Zixin Wen, Yuejie Chi +4
Reinforcement learning with verifiable rewards (RLVR) has been a main driver of recent breakthroughs in large reasoning models. Yet it remains a mystery how rewards based solely on…
Uncertainty quantification for Markov chain induced martingales with application to temporal difference learning
Weichen Wu, Yuting Wei, Alessandro Rinaldo
We establish novel and general high-dimensional concentration inequalities and Berry-Esseen bounds for vector-valued martingales induced by Markov chains. We apply these results to…
Statistical Inference under Adaptive Sampling with LinUCB
Wei Fan, Kevin Tan, Yuting Wei
Adaptively collected data has become ubiquitous within modern practice. However, even seemingly benign adaptive sampling schemes can introduce severe biases, rendering traditional…
Actor-Critics Can Achieve Optimal Sample Efficiency
Kevin Tan, Wei Fan, Yuting Wei
Actor-critic algorithms have become a cornerstone in reinforcement learning (RL), leveraging the strengths of both policy-based and value-based methods. Despite recent progress in…