1 paper
Michael K. Cohen, Marcus Hutter, Yoshua Bengio +1
In reinforcement learning, if the agent's reward differs from the designers' true utility, even only rarely, the state distribution resulting from the agent's policy can be very ba…