2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Ziqin Yuan, Ruiqi Wang, Dezhong Zhao +2
Preference-based reinforcement learning offers a scalable alternative to manual reward engineering by learning reward structures from comparative feedback. However, large-scale pre…