1 paper
Hao Ma, Shijie Wang, Zhiqiang Pu +2
Guiding the policy of multi-agent reinforcement learning to align with human common sense is a difficult problem, largely due to the complexity of modeling common sense as a reward…