activity
20122022
most citedEmphatic Temporal-Difference Learning

21 citations · 87 across the 18 of their papers we have counts for

collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI20204 cited

Maximizing Information Gain in Partially Observable Environments via Prediction Reward

Yash Satsangi, Sungsu Lim, Shimon Whiteson +2

Information gathering in a partially observable environment can be formulated as a reinforcement learning (RL), problem where the reward depends on the agent's uncertainty. For exa…

cs.AI2018

The Barbados 2018 List of Open Issues in Continual Learning

Tom Schaul, Hado van Hasselt, Joseph Modayil +7

We want to make progress toward artificial general intelligence, namely general-purpose agents that autonomously learn how to competently act in complex environments. The purpose o…

cs.AI2018

Organizing Experience: A Deeper Look at Replay Mechanisms for Sample-based Planning in Continuous State Domains

Yangchen Pan, Muhammad Zaheer, Adam White +2

Model-based strategies for control are critical to obtain sample efficient learning. Dyna is a planning paradigm that naturally interleaves learning and planning, by simulating one…

cs.AI2018

Directly Estimating the Variance of the λ-Return Using Temporal-Difference Methods

Craig Sherstan, Brendan Bennett, Kenny Young +4

This paper investigates estimating the variance of a temporal-difference learning agent's update target. Most reinforcement learning methods use an estimate of the value function,…

cs.AI20175 cited

Learning Sparse Representations in Reinforcement Learning with Sparse Coding

Lei Le, Raksha Kumaraswamy, Martha White

A variety of representation learning approaches have been investigated for reinforcement learning; much less attention, however, has been given to investigating the utility of spar…

cs.AI20179 cited

Data-Efficient Policy Evaluation Through Behavior Policy Search

Josiah P. Hanna, Philip S. Thomas, Peter Stone +1

We consider the task of evaluating a policy for a Markov decision process (MDP). The standard unbiased technique for evaluating a policy is to deploy the policy and observe its per…