activity
20122022
most citedGradient Temporal-Difference Learning with Regularized Corrections

9 citations · 25 across the 10 of their papers we have counts for

collaborators
Showing cs.AIShow all

8 papers · 1 filter

cs.AI20222 cited

What makes useful auxiliary tasks in reinforcement learning: investigating the effect of the target policy

Banafsheh Rafiee, Jun Jin, Jun Luo +1

Auxiliary tasks have been argued to be useful for representation learning in reinforcement learning. Although many auxiliary tasks have been empirically shown to be effective for a…

cs.AI20222 cited

The Frost Hollow Experiments: Pavlovian Signalling as a Path to Coordination and Communication Between Agents

Patrick M. Pilarski, Andrew Butcher, Elnaz Davoodi +7

Learned communication between agents is a powerful tool when approaching decision-making problems that are hard to overcome by any single agent in isolation. However, continual coo…

cs.AI20221 cited

Pavlovian Signalling with General Value Functions in Agent-Agent Temporal Decision Making

Andrew Butcher, Michael Bradley Johanson, Elnaz Davoodi +6

In this paper, we contribute a multi-faceted study into Pavlovian signalling -- a process by which learned, temporally extended predictions made by one agent inform decision-making…

cs.AI2018

The Barbados 2018 List of Open Issues in Continual Learning

Tom Schaul, Hado van Hasselt, Joseph Modayil +7

We want to make progress toward artificial general intelligence, namely general-purpose agents that autonomously learn how to competently act in complex environments. The purpose o…

cs.AI2018

Organizing Experience: A Deeper Look at Replay Mechanisms for Sample-based Planning in Continuous State Domains

Yangchen Pan, Muhammad Zaheer, Adam White +2

Model-based strategies for control are critical to obtain sample efficient learning. Dyna is a planning paradigm that naturally interleaves learning and planning, by simulating one…

cs.AI2018

Directly Estimating the Variance of the λ-Return Using Temporal-Difference Methods

Craig Sherstan, Brendan Bennett, Kenny Young +4

This paper investigates estimating the variance of a temporal-difference learning agent's update target. Most reinforcement learning methods use an estimate of the value function,…