1 paper
Viraj Mehta, Ojash Neopane, Vikramjeet Das +3
Preference-based feedback is important for many applications where direct evaluation of a reward function is not feasible. A notable recent example arises in reinforcement learning…