EPOpt: Learning Robust Neural Network Policies Using Model Ensembles
arXiv:1610.01283
Abstract
Sample complexity and safety are major challenges when learning policies with reinforcement learning for real-world tasks, especially when the policies are represented using rich function approximators like deep neural networks. Model-based methods where the real-world target domain is approximated using a simulated source domain provide an avenue to tackle the above challenges by augmenting real data with simulated data. However, discrepancies between the simulated source domain and the target domain pose a challenge for simulated training. We introduce the EPOpt algorithm, which uses an ensemble of simulated source domains and a form of adversarial training to learn policies that are robust and generalize to a broad range of possible target domains, including unmodeled effects. Further, the probability distribution over source domains in the ensemble can be adapted using data from target domain and approximate Bayesian methods, to progressively make it a better approximation. Thus, learning on a model ensemble, along with source domain adaptation, provides the benefit of both robustness and learning/adaptation.
Accepted for publication at the International Conference on Learning Representations (ICLR) 2017. Supplementary video: https://youtu.be/w1YJ9vwaoto
References in corpus (1)
Cited by in corpus (18)
- Robust Adversarial Reinforcement Learning
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
- Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
- Robust Deep Reinforcement Learning with Adversarial Attacks
- DeepRacer: Educational Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning
- Mutual Alignment Transfer Learning
- Locally Private Distributed Reinforcement Learning
- Deep Robust Kalman Filter
- Distributionally Robust Reinforcement Learning
- How You Act Tells a Lot: Privacy-Leakage Attack on Deep Reinforcement Learning
- Generalization through Simulation: Integrating Simulated and Real Data into Deep Reinforcement Learning for Vision-Based Autonomous Flight
- Map-based Multi-Policy Reinforcement Learning: Enhancing Adaptability of Robots by Deep Reinforcement Learning
- Learning Powerful Policies by Using Consistent Dynamics Model
- VMAV-C: A Deep Attention-based Reinforcement Learning Algorithm for Model-based Control
- Deep Model-Based Reinforcement Learning via Estimated Uncertainty and Conservative Policy Optimization
- Improved Robustness and Safety for Autonomous Vehicle Control with Adversarial Reinforcement Learning
- Quasi-Newton Trust Region Policy Optimization
- Motion Generation Considering Situation with Conditional Generative Adversarial Networks for Throwing Robots