Model-Ensemble Trust-Region Policy Optimization
arXiv:1802.10592
Abstract
Model-free reinforcement learning (RL) methods are succeeding in a growing number of tasks, aided by recent advances in deep learning. However, they tend to suffer from high sample complexity, which hinders their use in real-world domains. Alternatively, model-based reinforcement learning promises to reduce sample complexity, but tends to require careful tuning and to date have succeeded mainly in restrictive domains where simple models are sufficient for learning. In this paper, we analyze the behavior of vanilla model-based reinforcement learning methods when deep neural networks are used to learn both the model and the policy, and show that the learned policy tends to exploit regions where insufficient data is available for the model to be learned, causing instability in training. To overcome this issue, we propose to use an ensemble of models to maintain the model uncertainty and regularize the learning process. We further show that the use of likelihood ratio derivatives yields much more stable learning than backpropagation through time. Altogether, our approach Model-Ensemble Trust-Region Policy Optimization (ME-TRPO) significantly reduces the sample complexity compared to model-free deep RL methods on challenging continuous control benchmark tasks.
References in corpus (1)
Cited by in corpus (118)
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- Model-Based Reinforcement Learning for Atari
- MOPO: Model-based Offline Policy Optimization
- Reinforcement Learning with Augmented Data
- Benchmarking Model-Based Reinforcement Learning
- Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models
- Model-Based Value Estimation for Efficient Model-Free Reinforcement Learning
- MOReL : Model-Based Offline Reinforcement Learning
- Dream to Control: Learning Behaviors by Latent Imagination
- Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees
- When to Trust Your Model: Model-Based Policy Optimization
- Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning
- Exploring Model-based Planning with Policy Networks
- Dynamics-Aware Unsupervised Discovery of Skills
- Deep Dynamics Models for Learning Dexterous Manipulation
- Assessing Transferability from Simulation to Reality for Reinforcement Learning
- A Game Theoretic Framework for Model Based Reinforcement Learning
- A differentiable programming method for quantum control
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement Learning
- Effective Diversity in Population Based Reinforcement Learning
- Model-Augmented Actor-Critic: Backpropagating through Paths
- Search on the Replay Buffer: Bridging Planning and Reinforcement Learning
- Variational Inference MPC for Bayesian Model-based Reinforcement Learning
- Information Theoretic Regret Bounds for Online Nonlinear Control
- Model-based Reinforcement Learning for Semi-Markov Decision Processes with Neural ODEs
- Mastering Atari with Discrete World Models
- World Model as a Graph: Learning Latent Landmarks for Planning
- On-Policy Robot Imitation Learning from a Converging Supervisor
- Benchmarking deep inverse models over time, and the neural-adjoint method
- Robust Model-based Reinforcement Learning for Autonomous Greenhouse Control
- Dropout Q-Functions for Doubly Efficient Reinforcement Learning
- Trust the Model When It Is Confident: Masked Model-based Actor-Critic
- Learning to Manipulate Deformable Objects without Demonstrations
- Learning Robotic Manipulation through Visual Planning and Acting
- On the model-based stochastic value gradient for continuous reinforcement learning
- Model-based Adversarial Meta-Reinforcement Learning
- MBRL-Lib: A Modular Library for Model-based Reinforcement Learning
- Can Increasing Input Dimensionality Improve Deep Reinforcement Learning?
- Asynchronous Methods for Model-Based Reinforcement Learning
- Dynamics-Aware Quality-Diversity for Efficient Learning of Skill Repertoires
- Model-Predictive Control via Cross-Entropy and Gradient-Based Optimization
- Gradient-Aware Model-based Policy Search
- A Tutorial on Sparse Gaussian Processes and Variational Inference
- How to Learn a Useful Critic? Model-based Action-Gradient-Estimator Policy Optimization
- Reset-Free Lifelong Learning with Skill-Space Planning
- Safe Interactive Model-Based Learning
- Multi-Agent Reinforcement Learning with Multi-Step Generative Models
- Leveraging Digital Cousins for Ensemble Q-Learning in Large-Scale Wireless Networks
- Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment
- Model-based Lookahead Reinforcement Learning
- DyNODE: Neural Ordinary Differential Equations for Dynamics Modeling in Continuous Control
- Expert-Supervised Reinforcement Learning for Offline Policy Learning and Evaluation
- Uncertainty-aware Model-based Policy Optimization
- Floyd-Warshall Reinforcement Learning: Learning from Past Experiences to Reach New Goals
- Model-Based Opponent Modeling
- Autoregressive Dynamics Models for Offline Policy Evaluation and Optimization
- Direct and indirect reinforcement learning
- Efficient reinforcement learning control for continuum robots based on Inexplicit Prior Knowledge
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and Planning
- Adaptive Online Planning for Continual Lifelong Learning
- Policy-Aware Model Learning for Policy Gradient Methods
- Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative Sampling
- Deep Residual Reinforcement Learning
- MUSBO: Model-based Uncertainty Regularized and Sample Efficient Batch Optimization for Deployment Constrained Reinforcement Learning
- Information Theoretic Model Predictive Q-Learning
- MEPG: A Minimalist Ensemble Policy Gradient Framework for Deep Reinforcement Learning
- Regularizing Trajectory Optimization with Denoising Autoencoders
- Uncertainty-sensitive Learning and Planning with Ensembles
- Bridging Imagination and Reality for Model-Based Deep Reinforcement Learning
- Model-free and Bayesian Ensembling Model-based Deep Reinforcement Learning for Particle Accelerator Control Demonstrated on the FERMI FEL
- Untangling Dense Knots by Learning Task-Relevant Keypoints
- Selective Dyna-style Planning Under Limited Model Capacity
- Model Imitation for Model-Based Reinforcement Learning
- A Model-based Approach for Sample-efficient Multi-task Reinforcement Learning
- Sequence Generation with Guider Network
- Continuous-Time Model-Based Reinforcement Learning
- Learning Dynamics Models for Model Predictive Agents
- Few Shot System Identification for Reinforcement Learning
- On Effective Scheduling of Model-based Reinforcement Learning
- Model-Invariant State Abstractions for Model-Based Reinforcement Learning
- Control as Hybrid Inference
- Continuous Transition: Improving Sample Efficiency for Continuous Control Problems via MixUp
- A Fully Stochastic Second-Order Trust Region Method
- MobILE: Model-Based Imitation Learning From Observation Alone
- Centralized Model and Exploration Policy for Multi-Agent RL
- VMAV-C: A Deep Attention-based Reinforcement Learning Algorithm for Model-based Control
- Building Minimal and Reusable Causal State Abstractions for Reinforcement Learning
- PerSim: Data-Efficient Offline Reinforcement Learning with Heterogeneous Agents via Personalized Simulators
- Deep Model-Based Reinforcement Learning via Estimated Uncertainty and Conservative Policy Optimization
- CostNet: An End-to-End Framework for Goal-Directed Reinforcement Learning
- Baconian: A Unified Open-source Framework for Model-Based Reinforcement Learning
- Sample Efficient Reinforcement Learning via Model-Ensemble Exploration and Exploitation
- Policy Optimization with Model-based Explorations
- Imaginary Hindsight Experience Replay: Curious Model-based Learning for Sparse Reward Tasks
- A Free Lunch from the Noise: Provable and Practical Exploration for Representation Learning
- Physics-informed Dyna-Style Model-Based Deep Reinforcement Learning for Dynamic Control
- Learning Off-Policy with Online Planning
- An Active Learning Framework for Efficient Robust Policy Search
- Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles
- Continuous Control With Ensemble Deep Deterministic Policy Gradients
- On the Expressivity of Neural Networks for Deep Reinforcement Learning
- Mixed Reinforcement Learning with Additive Stochastic Uncertainty
- Optimal Control-Based Baseline for Guided Exploration in Policy Gradient Methods
- Learning to Reweight Imaginary Transitions for Model-Based Reinforcement Learning
- Improving Adversarial Text Generation by Modeling the Distant Future
- Model Embedding Model-Based Reinforcement Learning
- Model Predictive Actor-Critic: Accelerating Robot Skill Acquisition with Deep Reinforcement Learning
- MBDP: A Model-based Approach to Achieve both Robustness and Sample Efficiency via Double Dropout Planning
- Visual Reaction: Learning to Play Catch with Your Drone
- On-Policy Model Errors in Reinforcement Learning
- Distributed Ensembles of Reinforcement Learning Agents for Electricity Control
- Evaluating task-agnostic exploration for fixed-batch learning of arbitrary future tasks
- Eden: A Unified Environment Framework for Booming Reinforcement Learning Algorithms
- Continuous Deep Q-Learning with Simulator for Stabilization of Uncertain Discrete-Time Systems
- Using Human-Guided Causal Knowledge for More Generalized Robot Task Planning
- A Contraction Approach to Model-based Reinforcement Learning
- Bayes-Adaptive Deep Model-Based Policy Optimisation
- Planning with Exploration: Addressing Dynamics Bottleneck in Model-based Reinforcement Learning