12 papers
Logarithmic Regret for Online KL-Regularized Reinforcement Learning
Heyang Zhao, Chenlu Ye, Wei Xiong +2
Recent advances in Reinforcement Learning from Human Feedback (RLHF) have shown that KL-regularization plays a pivotal role in improving the efficiency of RL fine-tuning for large…
Towards a Sharp Analysis of Offline Policy Learning for -Divergence-Regularized Contextual Bandits
Qingyue Zhao, Kaixuan Ji, Heyang Zhao +2
Many offline reinforcement learning algorithms are underpinned by -divergence regularization, but their sample complexity *defined with respect to regularized objectives* still…
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
Qiwei Di, Jiafan He, Quanquan Gu
Learning from human feedback plays an important role in aligning generative models, such as large language models (LLM). However, the effectiveness of this approach can be influenc…
A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation
Heyang Zhao, Jiafan He, Quanquan Gu
The exploration-exploitation dilemma has been a central challenge in reinforcement learning (RL) with complex model classes. In this paper, we propose a new algorithm, Monotonic Q-…
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
Heyang Zhao, Xingrui Yu, David M. Bossens +2
Imitation learning is a central problem in reinforcement learning where the goal is to learn a policy that mimics the expert's behavior. In practice, it is often challenging to lea…
Variance-Dependent Regret Lower Bounds for Contextual Bandits
Jiafan He, Quanquan Gu
Variance-dependent regret bounds for linear contextual bandits, which improve upon the classical regret bound to , whe…