Virtual vs. Real: Trading Off Simulations and Physical Experiments in Reinforcement Learning with Bayesian Optimization
arXiv:1703.01250 · doi:10.1109/ICRA.2017.7989186
Abstract
In practice, the parameters of control policies are often tuned manually. This is time-consuming and frustrating. Reinforcement learning is a promising alternative that aims to automate this process, yet often requires too many experiments to be practical. In this paper, we propose a solution to this problem by exploiting prior knowledge from simulations, which are readily available for most robotic platforms. Specifically, we extend Entropy Search, a Bayesian optimization algorithm that maximizes information gain from each experiment, to the case of multiple information sources. The result is a principled way to automatically combine cheap, but inaccurate information from simulations with expensive and accurate physical experiments in a cost-effective manner. We apply the resulting method to a cart-pole system, which confirms that the algorithm can find good control policies with fewer experiments than standard Bayesian optimization on the physical system only.
7 pages, 6 figures, to appear in IEEE 2017 International Conference on Robotics and Automation (ICRA)
Cited by in corpus (32)
- A review of domain adaptation without target labels
- Data-efficient Auto-tuning with Bayesian Optimization: An Industrial Control Study
- Towards Dynamic and Safe Configuration Tuning for Cloud Databases
- Emulation of physical processes with Emukit
- A Search-Based Testing Approach for Deep Reinforcement Learning Agents
- Residual Policy Learning
- The Blind Implosion-Maker - Automated Inertial Confinement Fusion experiment design
- Derivative-Free Reinforcement Learning: A Review
- Resource-aware IoT Control: Saving Communication through Predictive Triggering
- On the Design of LQR Kernels for Efficient Controller Learning
- CybORG: An Autonomous Cyber Operations Research Gym
- Gait learning for soft microrobots controlled by light fields
- Mutual Alignment Transfer Learning
- Motion Planning by Reinforcement Learning for an Unmanned Aerial Vehicle in Virtual Open Space with Static Obstacles
- Learning a Low-dimensional Representation of a Safe Region for Safe Reinforcement Learning on Dynamical Systems
- Evaluating Low-Power Wireless Cyber-Physical Systems
- Bayesian Optimization for Policy Search via Online-Offline Experimentation
- Towards Assessing the Impact of Bayesian Optimization's Own Hyperparameters
- Identifying Mechanical Models through Differentiable Simulations
- Deep Kernels for Optimizing Locomotion Controllers
- Cautious Bayesian Optimization for Efficient and Scalable Policy Search
- A Multi-Fidelity Bayesian Approach to Safe Controller Design
- Computationally Efficient High-Dimensional Bayesian Optimization via Variable Selection
- Optimizing Photonic Nanostructures via Multi-fidelity Gaussian Processes
- Parameter Optimization using high-dimensional Bayesian Optimization
- Learning to Slide Unknown Objects with Differentiable Physics Simulations
- Knowledge Transfer Between Robots with Similar Dynamics for High-Accuracy Impromptu Trajectory Tracking
- NeoNav: Improving the Generalization of Visual Navigation via Generating Next Expected Observations
- Simulation-Aided Policy Tuning for Black-Box Robot Learning
- Multi-Fidelity Reinforcement Learning with Gaussian Processes
- Efficient Model Identification for Tensegrity Locomotion
- Enhancement of Energy-Based Swing-Up Controller via Entropy Search