7 papers
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
Qingyue Zhao, Kaixuan Ji, Heyang Zhao +1
\emph{Kullback-Leibler} (KL) regularization is ubiquitous in reinforcement learning algorithms in the form of \emph{reverse} or \emph{forward} KL. Recent studies have demonstrated…
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
Kaixuan Ji, Qiwei Di, Heyang Zhao +2
Kullback-Leibler (KL) regularization is widely used in offline decision-making and offers several benefits, motivating recent work on the sample complexity of offline learning with…
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
Kaixuan Ji, Qingyue Zhao, Heyang Zhao +2
Recent studies have shown that reinforcement learning with KL-regularized objectives can enjoy faster rates of convergence or logarithmic regret, in contrast to the classical $\sqr…
Towards a Sharp Analysis of Offline Policy Learning for -Divergence-Regularized Contextual Bandits
Qingyue Zhao, Kaixuan Ji, Heyang Zhao +2
Many offline reinforcement learning algorithms are underpinned by -divergence regularization, but their sample complexity *defined with respect to regularized objectives* still…
Best-of-Majority: Minimax-Optimal Strategy for Pass@ Inference Scaling
Qiwei Di, Kaixuan Ji, Xuheng Li +2
LLM inference often generates a batch of candidates for a prompt and selects one via strategies like majority voting or Best-of- N (BoN). For difficult tasks, this single-shot sele…
Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization
Kaixuan Ji, Guanlin Liu, Ning Dai +6
Reinforcement Learning (RL) plays a crucial role in aligning large language models (LLMs) with human preferences and improving their ability to perform complex tasks. However, curr…