38 citations · 41 across the 2 of their papers we have counts for
4 papers
Adaptive Temporal-Difference Learning for Policy Evaluation with Per-State Uncertainty Estimates
Hugo Penedones, Carlos Riquelme, Damien Vincent +5
We consider the core reinforcement-learning problem of on-policy value function approximation from a batch of trajectory data, and focus on various issues of Temporal Difference (T…
Fast Task Inference with Variational Intrinsic Successor Features
Steven Hansen, Will Dabney, Andre Barreto +3
It has been established that diverse behaviors spanning the controllable subspace of an Markov decision process can be trained by rewarding a policy for being distinguishable from…
Transfer in Deep Reinforcement Learning Using Successor Features and Generalised Policy Improvement
André Barreto, Diana Borsa, John Quan +6
The ability to transfer skills across tasks has the potential to scale up reinforcement learning (RL) agents to environments currently out of reach. Recently, a framework based on…
Composing Entropic Policies using Divergence Correction
Jonathan J Hunt, Andre Barreto, Timothy P Lillicrap +1
Composing previously mastered skills to solve novel tasks promises dramatic improvements in the data efficiency of reinforcement learning. Here, we analyze two recent works composi…