2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Xiaoyu Chen, Han Zhong, Zhuoran Yang +2
We study human-in-the-loop reinforcement learning (RL) with trajectory preferences, where instead of receiving a numeric reward at each step, the agent only receives preferences ov…