1 paper
Alekh Agarwal, Christoph Dann, Teodor V. Marinov
Offline algorithms for Reinforcement Learning from Human Preferences (RLHF), which use only a fixed dataset of sampled responses given an input, and preference feedback among these…