1 paper · 1 filter
Yulian Wu, Rushil Thareja, Praneeth Vepakomma +1
In this paper, we study the offline and online settings of reinforcement learning from human feedback (RLHF) with KL-regularization -- a widely used objective function in large lan…