1 paper
Xiao Hu, Jianxiong Li, Xianyuan Zhan +2
Preference-based reinforcement learning (PbRL) provides a natural way to align RL agents' behavior with human desired outcomes, but is often restrained by costly human feedback. To…