3 papers
cs.LG2019
Incrementally Learning Functions of the Return
Brendan Bennett, Wesley Chung, Muhammad Zaheer +1
Temporal difference methods enable efficient estimation of value functions in reinforcement learning in an incremental fashion, and are of broader interest because they correspond…
cs.LG2019
Planning with Expectation Models
Yi Wan, Zaheer Abbas, Adam White +2
Distribution and sample models are two popular model choices in model-based reinforcement learning (MBRL). However, learning these models can be intractable, particularly when the…
cs.AI2018
Organizing Experience: A Deeper Look at Replay Mechanisms for Sample-based Planning in Continuous State Domains
Yangchen Pan, Muhammad Zaheer, Adam White +2
Model-based strategies for control are critical to obtain sample efficient learning. Dyna is a planning paradigm that naturally interleaves learning and planning, by simulating one…