3 citations · 6 across the 5 of their papers we have counts for
1 paper · 1 filter
Michael K. Cohen, Marcus Hutter, Yoshua Bengio +1
In reinforcement learning, if the agent's reward differs from the designers' true utility, even only rarely, the state distribution resulting from the agent's policy can be very ba…