Self-Correcting Models for Model-Based Reinforcement Learning
arXiv:1612.06018
Abstract
When an agent cannot represent a perfectly accurate model of its environment's dynamics, model-based reinforcement learning (MBRL) can fail catastrophically. Planning involves composing the predictions of the model; when flawed predictions are composed, even minor errors can compound and render the model useless for planning. Hallucinated Replay (Talvitie 2014) trains the model to "correct" itself when it produces errors, substantially improving MBRL with flawed models. This paper theoretically analyzes this approach, illuminates settings in which it is likely to be effective or ineffective, and presents a novel error bound, showing that a model's ability to self-correct is more tightly related to MBRL performance than one-step prediction error. These results inspire an MBRL algorithm for deterministic MDPs with performance guarantees that are robust to model class limitations.
Original paper appeared in Proceedings of the 31st AAAI Conference on Artificial Intelligence, 2017. This version incorporates the appendix into document (rather than as supplementary material), corrects a minor error in Lemma 1, and fixes some type-os
Cited by in corpus (17)
- When to Trust Your Model: Model-Based Policy Optimization
- MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments
- Data Efficient Reinforcement Learning for Legged Robots
- Combating the Compounding-Error Problem with a Multi-step Model
- Objective Mismatch in Model-based Reinforcement Learning
- Learning to Combat Compounding-Error in Model-Based Reinforcement Learning
- Deep Residual Reinforcement Learning
- Policy-Aware Model Learning for Policy Gradient Methods
- Model Imitation for Model-Based Reinforcement Learning
- Selective Dyna-style Planning Under Limited Model Capacity
- Model Primitive Hierarchical Lifelong Reinforcement Learning
- Frequency-based Search-control in Dyna
- Temporally Abstract Partial Models
- Planning with Expectation Models for Control
- On-Policy Model Errors in Reinforcement Learning
- Domain Knowledge Integration By Gradient Matching For Sample-Efficient Reinforcement Learning
- Dynamic Horizon Value Estimation for Model-based Reinforcement Learning