2 papers
cs.LG2026
Efficient Exploration at Scale
Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla +5
We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF). Our algorithm incrementally updates rew…
cs.LG2023
Efficient Online Learning with Offline Datasets for Infinite Horizon MDPs: A Bayesian Approach
Dengwang Tang, Rahul Jain, Botao Hao +1
In this paper, we study the problem of efficient online reinforcement learning in the infinite horizon setting when there is an offline dataset to start with. We assume that the of…