1 paper
Zehong Cao, KaiChiu Wong, Chin-Teng Lin
The current reward learning from human preferences could be used to resolve complex reinforcement learning (RL) tasks without access to a reward function by defining a single fixed…