3 citations · 4 across the 2 of their papers we have counts for
1 paper · 1 filter
Yichong Xu, Ruosong Wang, Lin F. Yang +2
Preference-based Reinforcement Learning (PbRL) replaces reward values in traditional reinforcement learning by preferences to better elicit human opinion on the target objective, e…