30 citations · 66 across the 24 of their papers we have counts for
25 papers · 1 filter
Variance-Dependent Regret Lower Bounds for Contextual Bandits
Jiafan He, Quanquan Gu
Variance-dependent regret bounds for linear contextual bandits, which improve upon the classical regret bound to , wher…
Accelerated Preference Optimization for Large Language Model Alignment
Jiafan He, Huizhuo Yuan, Quanquan Gu
Reinforcement Learning from Human Feedback (RLHF) has emerged as a pivotal tool for aligning large language models (LLMs) with human preferences. Direct Preference Optimization (DP…
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
Qiwei Di, Jiafan He, Quanquan Gu
Learning from human feedback plays an important role in aligning generative models, such as large language models (LLM). However, the effectiveness of this approach can be influenc…
Achieving Constant Regret in Linear Markov Decision Processes
Weitong Zhang, Zhiyuan Fan, Jiafan He +1
We study the constant regret guarantees in reinforcement learning (RL). Our objective is to design an algorithm that incurs only finite regret over infinite episodes with high prob…
Nearly Minimax Optimal Regret for Learning Linear Mixture Stochastic Shortest Path
Qiwei Di, Jiafan He, Dongruo Zhou +1
We study the Stochastic Shortest Path (SSP) problem with a linear mixture transition kernel, where an agent repeatedly interacts with a stochastic environment and seeks to reach ce…
Reinforcement Learning from Human Feedback with Active Queries
Kaixuan Ji, Jiafan He, Quanquan Gu
Aligning large language models (LLM) with human preference plays a key role in building modern generative models and can be achieved by reinforcement learning from human feedback (…