Model-Based Value Estimation for Efficient Model-Free Reinforcement Learning
arXiv:1803.00101
Abstract
Recent model-free reinforcement learning algorithms have proposed incorporating learned dynamics models as a source of additional data with the intention of reducing sample complexity. Such methods hold the promise of incorporating imagined data coupled with a notion of model uncertainty to accelerate the learning of continuous control tasks. Unfortunately, they rely on heuristics that limit usage of the dynamics model. We present model-based value expansion, which controls for uncertainty in the model by only allowing imagination to fixed depth. By enabling wider use of learned dynamics models within a model-free reinforcement learning algorithm, we improve value estimation, which, in turn, reduces the sample complexity of learning.
References in corpus (4)
Cited by in corpus (76)
- Model-Based Reinforcement Learning for Atari
- Dream to Control: Learning Behaviors by Latent Imagination
- Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees
- When to Trust Your Model: Model-Based Policy Optimization
- A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms
- Exploring Model-based Planning with Policy Networks
- DRLinFluids -- An open-source python platform of coupling Deep Reinforcement Learning and OpenFOAM
- DeepMDP: Learning Continuous Latent Space Models for Representation Learning
- A Survey on Reinforcement Learning in Aviation Applications
- A Game Theoretic Framework for Model Based Reinforcement Learning
- Softmax Deep Double Deterministic Policy Gradients
- Off-Policy Multi-Agent Decomposed Policy Gradients
- Model-Augmented Actor-Critic: Backpropagating through Paths
- Learning to Fly via Deep Model-Based Reinforcement Learning
- Rethinking Closed-loop Training for Autonomous Driving
- Sample-efficient Reinforcement Learning Representation Learning with Curiosity Contrastive Forward Dynamics Model
- Trust the Model When It Is Confident: Masked Model-based Actor-Critic
- Improving Robot Dual-System Motor Learning with Intrinsically Motivated Meta-Control and Latent-Space Experience Imagination
- Learning to Combat Compounding-Error in Model-Based Reinforcement Learning
- Sample-Efficient Reinforcement Learning via Counterfactual-Based Data Augmentation
- Probabilistic design of optimal sequential decision-making algorithms in learning and control
- Model-based Adversarial Meta-Reinforcement Learning
- Reward Estimation for Variance Reduction in Deep Reinforcement Learning
- Gradient-Aware Model-based Policy Search
- Combining Model-Based and Model-Free Methods for Nonlinear Control: A Provably Convergent Policy Gradient Approach
- How to Learn a Useful Critic? Model-based Action-Gradient-Estimator Policy Optimization
- Task-Agnostic Dynamics Priors for Deep Reinforcement Learning
- DyNODE: Neural Ordinary Differential Equations for Dynamics Modeling in Continuous Control
- Model-Based Opponent Modeling
- Model-based Lookahead Reinforcement Learning
- The Effect of Multi-step Methods on Overestimation in Deep Reinforcement Learning
- Ready Policy One: World Building Through Active Learning
- Value Iteration in Continuous Actions, States and Time
- Efficient reinforcement learning control for continuum robots based on Inexplicit Prior Knowledge
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and Planning
- Direct and indirect reinforcement learning
- Deep Residual Reinforcement Learning
- Imagined Value Gradients: Model-Based Policy Optimization with Transferable Latent Dynamics Models
- Model-based Policy Optimization with Unsupervised Model Adaptation
- Bridging Imagination and Reality for Model-Based Deep Reinforcement Learning
- Efficient and Robust Reinforcement Learning with Uncertainty-based Value Expansion
- MUSBO: Model-based Uncertainty Regularized and Sample Efficient Batch Optimization for Deployment Constrained Reinforcement Learning
- Influence-Based Multi-Agent Exploration
- Selective Dyna-style Planning Under Limited Model Capacity
- Visual Adversarial Imitation Learning using Variational Models
- Human-in-the-Loop Methods for Data-Driven and Reinforcement Learning Systems
- Learning Dynamics Models for Model Predictive Agents
- On Effective Scheduling of Model-based Reinforcement Learning
- Uncertainty-Aware Model-Based Reinforcement Learning with Application to Autonomous Driving
- Towards a Simple Approach to Multi-step Model-based Reinforcement Learning
- Bias-reduced Multi-step Hindsight Experience Replay for Efficient Multi-goal Reinforcement Learning
- Learning Trajectories for Visual-Inertial System Calibration via Model-based Heuristic Deep Reinforcement Learning
- Building Minimal and Reusable Causal State Abstractions for Reinforcement Learning
- Imaginary Hindsight Experience Replay: Curious Model-based Learning for Sparse Reward Tasks
- Maximum Entropy Model Rollouts: Fast Model Based Policy Optimization without Compounding Errors
- MapGo: Model-Assisted Policy Optimization for Goal-Oriented Tasks
- TDprop: Does Jacobi Preconditioning Help Temporal Difference Learning?
- Sample Efficient Reinforcement Learning via Model-Ensemble Exploration and Exploitation
- Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles
- Learning Accurate Extended-Horizon Predictions of High Dimensional Trajectories
- Learning Off-Policy with Online Planning
- Improve Agents without Retraining: Parallel Tree Search with Off-Policy Correction
- On-Policy Model Errors in Reinforcement Learning
- A selected review on reinforcement learning based control for autonomous underwater vehicles
- Dynamic Horizon Value Estimation for Model-based Reinforcement Learning
- Planning with Expectation Models for Control
- Discriminator Augmented Model-Based Reinforcement Learning
- Low-Variance Policy Gradient Estimation with World Models
- Deep Reinforcement Learning with Adjustments
- Eden: A Unified Environment Framework for Booming Reinforcement Learning Algorithms
- Stochastic Intervention for Causal Inference via Reinforcement Learning
- Using Time-Series Privileged Information for Provably Efficient Learning of Prediction Models
- Generative Temporal Difference Learning for Infinite-Horizon Prediction
- Planning with Exploration: Addressing Dynamics Bottleneck in Model-based Reinforcement Learning
- Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
- When Autonomous Systems Meet Accuracy and Transferability through AI: A Survey