1 paper
Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla +5
We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF). Our algorithm incrementally updates rew…