21 citations · 87 across the 18 of their papers we have counts for
6 papers · 1 filter
Maximizing Information Gain in Partially Observable Environments via Prediction Reward
Yash Satsangi, Sungsu Lim, Shimon Whiteson +2
Information gathering in a partially observable environment can be formulated as a reinforcement learning (RL), problem where the reward depends on the agent's uncertainty. For exa…
The Barbados 2018 List of Open Issues in Continual Learning
Tom Schaul, Hado van Hasselt, Joseph Modayil +7
We want to make progress toward artificial general intelligence, namely general-purpose agents that autonomously learn how to competently act in complex environments. The purpose o…
Organizing Experience: A Deeper Look at Replay Mechanisms for Sample-based Planning in Continuous State Domains
Yangchen Pan, Muhammad Zaheer, Adam White +2
Model-based strategies for control are critical to obtain sample efficient learning. Dyna is a planning paradigm that naturally interleaves learning and planning, by simulating one…
Directly Estimating the Variance of the λ-Return Using Temporal-Difference Methods
Craig Sherstan, Brendan Bennett, Kenny Young +4
This paper investigates estimating the variance of a temporal-difference learning agent's update target. Most reinforcement learning methods use an estimate of the value function,…
Learning Sparse Representations in Reinforcement Learning with Sparse Coding
Lei Le, Raksha Kumaraswamy, Martha White
A variety of representation learning approaches have been investigated for reinforcement learning; much less attention, however, has been given to investigating the utility of spar…
Data-Efficient Policy Evaluation Through Behavior Policy Search
Josiah P. Hanna, Philip S. Thomas, Peter Stone +1
We consider the task of evaluating a policy for a Markov decision process (MDP). The standard unbiased technique for evaluating a policy is to deploy the policy and observe its per…