2 citations · 2 across the 1 of their papers we have counts for
1 paper
Peter Barnett, Rachel Freedman, Justin Svegliato +1
Reward learning algorithms utilize human feedback to infer a reward function, which is then used to train an AI system. This human feedback is often a preference comparison, in whi…