1 paper
Yijun Shen, Delong Chen, Xianming Hu +4
Agents must infer action outcomes and select actions that maximize a reward signal indicating how close the goal is to being reached. Supervised learning of reward models could int…