activity
20122022
most citedRainbow: Combining Improvements in Deep Reinforcement Learning

424 citations · 492 across the 7 of their papers we have counts for

collaborators

8 papers

cs.LG2021

Adapting the Function Approximation Architecture in Online Reinforcement Learning

John D. Martin, Joseph Modayil

The performance of a reinforcement learning (RL) system depends on the computational architecture used to approximate a value function. Deep learning methods provide both optimizat…

cs.LG201921 cited

On Inductive Biases in Deep Reinforcement Learning

Matteo Hessel, Hado van Hasselt, Joseph Modayil +1

Many deep reinforcement learning algorithms contain inductive biases that sculpt the agent's objective and its interface to the environment. These inductive biases can take many fo…

cs.LG201939 cited

Ray Interference: a Source of Plateaus in Deep Reinforcement Learning

Tom Schaul, Diana Borsa, Joseph Modayil +1

Rather than proposing a new method, this paper investigates an issue present in existing learning algorithms. We study the learning dynamics of reinforcement learning (RL), specifi…

cs.AI2018

Deep Reinforcement Learning and the Deadly Triad

Hado van Hasselt, Yotam Doron, Florian Strub +3

We know from reinforcement learning theory that temporal difference learning can fail in certain cases. Sutton and Barto (2018) identify a deadly triad of function approximation, b…

cs.AI2018

The Barbados 2018 List of Open Issues in Continual Learning

Tom Schaul, Hado van Hasselt, Joseph Modayil +7

We want to make progress toward artificial general intelligence, namely general-purpose agents that autonomously learn how to competently act in complex environments. The purpose o…

cs.AI20175 cited

Building Machines that Learn and Think for Themselves: Commentary on Lake et al., Behavioral and Brain Sciences, 2017

M. Botvinick, D. G. T. Barrett, P. Battaglia +16

We agree with Lake and colleagues on their list of key ingredients for building humanlike intelligence, including the idea that model-based reasoning is essential. However, we favo…