6 papers
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
Kaixuan Ji, Qiwei Di, Heyang Zhao +2
Kullback-Leibler (KL) regularization is widely used in offline decision-making and offers several benefits, motivating recent work on the sample complexity of offline learning with…
Dimension-Independent Convergence of Underdamped Langevin Monte Carlo in KL Divergence
Shiyuan Zhang, Qiwei Di, Xuheng Li +1
Underdamped Langevin dynamics (ULD) is a widely-used sampler for Gibbs distributions , and is often empirically effective in high dimensions. However, existing non…
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
Kaixuan Ji, Qingyue Zhao, Heyang Zhao +2
Recent studies have shown that reinforcement learning with KL-regularized objectives can enjoy faster rates of convergence or logarithmic regret, in contrast to the classical $\sqr…
On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
Yue Yu, Qiwei Di, Quanquan Gu +1
Test-time compute (TTC) has become an increasingly prominent paradigm for enhancing large language models (LLMs). Despite the empirical success of methods such as best-of- (BoN)…
Best-of-Majority: Minimax-Optimal Strategy for Pass@ Inference Scaling
Qiwei Di, Kaixuan Ji, Xuheng Li +2
LLM inference often generates a batch of candidates for a prompt and selects one via strategies like majority voting or Best-of- N (BoN). For difficult tasks, this single-shot sele…
Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers
Runjia Li, Qiwei Di, Quanquan Gu
Score-based diffusion models have emerged as powerful techniques for generating samples from high-dimensional data distributions. These models involve a two-phase process: first, i…