The Predictron: End-To-End Learning and Planning
arXiv:1612.08810
Abstract
One of the key challenges of artificial intelligence is to learn models that are effective in the context of planning. In this document we introduce the predictron architecture. The predictron consists of a fully abstract model, represented by a Markov reward process, that can be rolled forward multiple "imagined" planning steps. Each forward pass of the predictron accumulates internal rewards and values over multiple planning depths. The predictron is trained end-to-end so as to make these accumulated values accurately approximate the true value function. We applied the predictron to procedurally generated random mazes and a simulator for the game of pool. The predictron yielded significantly more accurate predictions than conventional deep neural network architectures.
Camera-ready version, ICML 2017, with supplement
References in corpus (1)
Cited by in corpus (60)
- An Introduction to Deep Reinforcement Learning
- Deep Reinforcement Learning: An Overview
- Learning to reinforcement learn
- Go-Explore: a New Approach for Hard-Exploration Problems
- Differentiable MPC for End-to-end Planning and Control
- A Review of Tracking, Prediction and Decision Making Methods for Autonomous Driving
- Learning Particle Dynamics for Manipulating Rigid Bodies, Deformable Objects, and Fluids
- When to Trust Your Model: Model-Based Policy Optimization
- Universal Planning Networks
- DeepMDP: Learning Continuous Latent Space Models for Representation Learning
- Learning and Querying Fast Generative Models for Reinforcement Learning
- Value Prediction Network
- Combating the Compounding-Error Problem with a Multi-step Model
- Learning to Search with MCTSnets
- Look Before You Leap: Bridging Model-Free and Model-Based Reinforcement Learning for Planned-Ahead Vision-and-Language Navigation
- Combining Q-Learning and Search with Amortized Value Estimates
- Shaping Belief States with Generative Environment Models for RL
- Value Propagation Networks
- Policy Evaluation Networks
- Gradient-Aware Model-based Policy Search
- Reward Estimation for Variance Reduction in Deep Reinforcement Learning
- Propagation Networks for Model-Based Control Under Partial Observation
- The Value Equivalence Principle for Model-Based Reinforcement Learning
- Discriminative Deep Dyna-Q: Robust Planning for Dialogue Policy Learning
- Learning to Predict Without Looking Ahead: World Models Without Forward Prediction
- Adapting Behaviour via Intrinsic Reward: A Survey and Empirical Study
- Relational Graph Learning for Crowd Navigation
- Deep Residual Reinforcement Learning
- Policy-Aware Model Learning for Policy Gradient Methods
- Deep Model-Based Reinforcement Learning for High-Dimensional Problems, a Survey
- XLVIN: eXecuted Latent Value Iteration Nets
- Adaptive Skip Intervals: Temporal Abstraction for Recurrent Dynamical Models
- Probing Emergent Semantics in Predictive Agents via Question Answering
- Steady State Analysis of Episodic Reinforcement Learning
- Gated Path Planning Networks
- Importance Resampling for Off-policy Prediction
- Document-editing Assistants and Model-based Reinforcement Learning as a Path to Conversational AI
- Towards a Simple Approach to Multi-step Model-based Reinforcement Learning
- Forethought and Hindsight in Credit Assignment
- Planning from Pixels using Inverse Dynamics Models
- Inductive-bias-driven Reinforcement Learning For Efficient Schedules in Heterogeneous Clusters
- Frequency-based Search-control in Dyna
- Beyond Tabula-Rasa: a Modular Reinforcement Learning Approach for Physically Embedded 3D Sokoban
- Value-driven Hindsight Modelling
- The Barbados 2018 List of Open Issues in Continual Learning
- Sufficiently Accurate Model Learning
- On the Expressivity of Neural Networks for Deep Reinforcement Learning
- Critical Percolation as a Framework to Analyze the Training of Deep Networks
- NeoNav: Improving the Generalization of Visual Navigation via Generating Next Expected Observations
- Amanuensis: The Programmer's Apprentice
- An investigation of model-free planning
- Visual Reaction: Learning to Play Catch with Your Drone
- Planning in Dynamic Environments with Conditional Autoregressive Models
- Generative Temporal Difference Learning for Infinite-Horizon Prediction
- Combining Off and On-Policy Training in Model-Based Reinforcement Learning
- Time-Aware Q-Networks: Resolving Temporal Irregularity for Deep Reinforcement Learning
- Visual Goal-Directed Meta-Learning with Contextual Planning Networks
- Differentiable Robust LQR Layers
- Neural Algorithmic Reasoners are Implicit Planners
- Transfer Value Iteration Networks