activity
20242026
collaborators

6 papers

cs.LG2026

On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization

Kaixuan Ji, Qiwei Di, Heyang Zhao +2

Kullback-Leibler (KL) regularization is widely used in offline decision-making and offers several benefits, motivating recent work on the sample complexity of offline learning with…

cs.LG2026

Dimension-Independent Convergence of Underdamped Langevin Monte Carlo in KL Divergence

Shiyuan Zhang, Qiwei Di, Xuheng Li +1

Underdamped Langevin dynamics (ULD) is a widely-used sampler for Gibbs distributions , and is often empirically effective in high dimensions. However, existing non…

cs.LG2026

Near-Optimal Regret for KL-Regularized Multi-Armed Bandits

Kaixuan Ji, Qingyue Zhao, Heyang Zhao +2

Recent studies have shown that reinforcement learning with KL-regularized objectives can enjoy faster rates of convergence or logarithmic regret, in contrast to the classical $\sqr…

cs.LG2025

On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference

Yue Yu, Qiwei Di, Quanquan Gu +1

Test-time compute (TTC) has become an increasingly prominent paradigm for enhancing large language models (LLMs). Despite the empirical success of methods such as best-of- (BoN)…

cs.LG2025

Best-of-Majority: Minimax-Optimal Strategy for Pass@ Inference Scaling

Qiwei Di, Kaixuan Ji, Xuheng Li +2

LLM inference often generates a batch of candidates for a prompt and selects one via strategies like majority voting or Best-of- N (BoN). For difficult tasks, this single-shot sele…

cs.LG2024

Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers

Runjia Li, Qiwei Di, Quanquan Gu

Score-based diffusion models have emerged as powerful techniques for generating samples from high-dimensional data distributions. These models involve a two-phase process: first, i…