1 paper · 1 filter
Luise Ge, Daniel Halpern, Evi Micha +4
In the context of reinforcement learning from human feedback (RLHF), the reward function is generally derived from maximum likelihood estimation of a random utility model based on…