Dyna-Style Planning with Linear Function Approximation and Prioritized Sweeping
arXiv:1206.3285
Abstract
We consider the problem of efficiently learning optimal control policies and value functions over large state spaces in an online setting in which estimates must be available after each interaction with the world. This paper develops an explicitly model-based approach extending the Dyna architecture to linear function approximation. Dynastyle planning proceeds by generating imaginary experience from the world model and then applying model-free reinforcement learning algorithms to the imagined state transitions. Our main results are to prove that linear Dyna-style planning converges to a unique solution independent of the generating distribution, under natural conditions. In the policy evaluation setting, we prove that the limit point is the least-squares (LSTD) solution. An implication of our results is that prioritized-sweeping can be soundly extended to the linear approximation case, backing up to preceding features rather than to preceding states. We introduce two versions of prioritized sweeping with linear Dyna and briefly illustrate their performance empirically on the Mountain Car and Boyan Chain problems.
Appears in Proceedings of the Twenty-Fourth Conference on Uncertainty in Artificial Intelligence (UAI2008)
Cited by in corpus (25)
- DualDICE: Behavior-Agnostic Estimation of Discounted Stationary Distribution Corrections
- Projective simulation for classical learning agents: a comprehensive investigation
- Value Prediction Network
- Towards a Common Implementation of Reinforcement Learning for Multiple Robotic Tasks
- Lipschitz Continuity in Model-based Reinforcement Learning
- Combating the Compounding-Error Problem with a Multi-step Model
- CoinDICE: Off-Policy Confidence Interval Estimation
- The Value Equivalence Principle for Model-Based Reinforcement Learning
- Organizing Experience: A Deeper Look at Replay Mechanisms for Sample-based Planning in Continuous State Domains
- Model-based Lookahead Reinforcement Learning
- Policy-Aware Model Learning for Policy Gradient Methods
- Model-Based Regularization for Deep Reinforcement Learning with Transcoder Networks
- The Effects of Memory Replay in Reinforcement Learning
- Document-editing Assistants and Model-based Reinforcement Learning as a Path to Conversational AI
- Forethought and Hindsight in Credit Assignment
- Towards a Simple Approach to Multi-step Model-based Reinforcement Learning
- Dyna Planning using a Feature Based Generative Model
- Affordance as general value function: A computational model
- Fuzzy Tiling Activations: A Simple Approach to Learning Sparse Representations Online
- Frequency-based Search-control in Dyna
- Projective simulation for artificial intelligence
- Hill Climbing on Value Estimates for Search-control in Dyna
- Planning with Expectation Models for Control
- Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
- Optimistic Simulated Exploration as an Incentive for Real Exploration