1 paper · 1 filter
Yijun Shen, Delong Chen, Xianming Hu +4
Agents must infer action outcomes and select actions that maximize a reward signal indicating how close the goal is to being reached. Supervised learning of reward models could int…