Planning from Pixels using Inverse Dynamics Models
arXiv:2012.02419
Abstract
Learning task-agnostic dynamics models in high-dimensional observation spaces can be challenging for model-based RL agents. We propose a novel way to learn latent world models by learning to predict sequences of future actions conditioned on task completion. These task-conditioned models adaptively focus modeling capacity on task-relevant dynamics, while simultaneously serving as an effective heuristic for planning with sparse rewards. We evaluate our method on challenging visual goal completion tasks and show a substantial increase in performance compared to prior model-free approaches.
9 pages, 4 figures
References in corpus (5)
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Rainbow: Combining Improvements in Deep Reinforcement Learning
- Benchmarking Model-Based Reinforcement Learning
- rlpyt: A Research Code Base for Deep Reinforcement Learning in PyTorch
- Causally Correct Partial Models for Reinforcement Learning