39 citations · 168 across the 9 of their papers we have counts for
14 papers
The Option Keyboard: Combining Skills in Reinforcement Learning
André Barreto, Diana Borsa, Shaobo Hou +8
The ability to combine known skills to create new ones may be crucial in the solution of complex reinforcement learning problems that unfold over extended periods. We argue that a…
Return-based Scaling: Yet Another Normalisation Trick for Deep RL
Tom Schaul, Georg Ostrovski, Iurii Kemaev +1
Scaling issues are mundane yet irritating for practitioners of reinforcement learning. Error scales vary across domains, tasks, and stages of learning; sometimes by many orders of…
Temporal Difference Uncertainties as a Signal for Exploration
Sebastian Flennerhag, Jane X. Wang, Pablo Sprechmann +7
An effective approach to exploration in reinforcement learning is to rely on an agent's uncertainty over the optimal policy, which can yield near-optimal exploration strategies in…
Expected Eligibility Traces
Hado van Hasselt, Sephora Madjiheurem, Matteo Hessel +3
The question of how to determine which states and actions are responsible for a certain outcome is known as the credit assignment problem and remains a central research question in…
Adapting Behaviour for Learning Progress
Tom Schaul, Diana Borsa, David Ding +4
Determining what experience to generate to best facilitate learning (i.e. exploration) is one of the distinguishing features and open challenges in reinforcement learning. The adve…
Conditional Importance Sampling for Off-Policy Learning
Mark Rowland, Anna Harutyunyan, Hado van Hasselt +4
The principal contribution of this paper is a conceptual framework for off-policy reinforcement learning, based on conditional expectations of importance sampling ratios. This fram…