30 citations · 89 across the 8 of their papers we have counts for
1 paper · 2 filters
Daniel Shin, Daniel S. Brown, Anca D. Dragan
Learning a reward function from human preferences is challenging as it typically requires having a high-fidelity simulator or using expensive and potentially unsafe actual physical…