153 citations · 273 across the 25 of their papers we have counts for
4 papers · 2 filters
Proper Value Equivalence
Christopher Grimm, André Barreto, Gregory Farquhar +2
One of the main challenges in model-based reinforcement learning (RL) is to decide which aspects of the environment should be modeled. The value-equivalence (VE) principle proposes…
Discovering Diverse Nearly Optimal Policies with Successor Features
Tom Zahavy, Brendan O'Donoghue, Andre Barreto +3
Finding different solutions to the same problem is a key aspect of intelligence associated with creativity and adaptation to novel situations. In reinforcement learning, a set of d…
Reward is enough for convex MDPs
Tom Zahavy, Brendan O'Donoghue, Guillaume Desjardins +1
Maximising a cumulative reward function that is Markov and stationary, i.e., defined over state-action pairs and independent of time, is sufficient to capture many kinds of goals i…
Discovering a set of policies for the worst case reward
Tom Zahavy, Andre Barreto, Daniel J Mankowitz +4
We study the problem of how to construct a set of policies that can be composed together to solve a collection of reinforcement learning tasks. Each task is a different reward func…