1 paper
Zibin Dong, Yifu Yuan, Jianye Hao +7
Aligning agent behaviors with diverse human preferences remains a challenging problem in reinforcement learning (RL), owing to the inherent abstractness and mutability of human pre…