collaborators

7 papers

cs.LG2026

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability

Qingyue Zhao, Kaixuan Ji, Heyang Zhao +1

\emph{Kullback-Leibler} (KL) regularization is ubiquitous in reinforcement learning algorithms in the form of \emph{reverse} or \emph{forward} KL. Recent studies have demonstrated…

cs.LG2026

On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization

Kaixuan Ji, Qiwei Di, Heyang Zhao +2

Kullback-Leibler (KL) regularization is widely used in offline decision-making and offers several benefits, motivating recent work on the sample complexity of offline learning with…

cs.LG2026

Near-Optimal Regret for KL-Regularized Multi-Armed Bandits

Kaixuan Ji, Qingyue Zhao, Heyang Zhao +2

Recent studies have shown that reinforcement learning with KL-regularized objectives can enjoy faster rates of convergence or logarithmic regret, in contrast to the classical $\sqr…

cs.LG2026

Towards a Sharp Analysis of Offline Policy Learning for -Divergence-Regularized Contextual Bandits

Qingyue Zhao, Kaixuan Ji, Heyang Zhao +2

Many offline reinforcement learning algorithms are underpinned by -divergence regularization, but their sample complexity *defined with respect to regularized objectives* still…

cs.LG2025

Best-of-Majority: Minimax-Optimal Strategy for Pass@ Inference Scaling

Qiwei Di, Kaixuan Ji, Xuheng Li +2

LLM inference often generates a batch of candidates for a prompt and selects one via strategies like majority voting or Best-of- N (BoN). For difficult tasks, this single-shot sele…

cs.LG2025

Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

Kaixuan Ji, Guanlin Liu, Ning Dai +6

Reinforcement Learning (RL) plays a crucial role in aligning large language models (LLMs) with human preferences and improving their ability to perform complex tasks. However, curr…