18 citations · 50 across the 17 of their papers we have counts for
1 paper · 2 filters
Mudit Verma, Katherine Metcalf
Specifying rewards for reinforcement learned (RL) agents is challenging. Preference-based RL (PbRL) mitigates these challenges by inferring a reward from feedback over sets of traj…