3 papers
cs.LG2018
Online Off-policy Prediction
Sina Ghiassian, Andrew Patterson, Martha White +2
This paper investigates the problem of online prediction learning, where learning proceeds continuously as the agent interacts with an environment. The predictions made by the agen…
cs.LG2018
General Value Function Networks
Matthew Schlegel, Andrew Jacobsen, Zaheer Abbas +3
State construction is important for learning in partially observable environments. A general purpose strategy for state construction is to learn the state update using a Recurrent…
cs.AI2018
Organizing Experience: A Deeper Look at Replay Mechanisms for Sample-based Planning in Continuous State Domains
Yangchen Pan, Muhammad Zaheer, Adam White +2
Model-based strategies for control are critical to obtain sample efficient learning. Dyna is a planning paradigm that naturally interleaves learning and planning, by simulating one…