1 paper
Eric Eaton, Marcel Hussing, Michael Kearns +3
In traditional reinforcement learning (RL), the learner aims to solve a single objective optimization problem: find the policy that maximizes expected reward. However, in many real…