Safe Exploration in Finite Markov Decision Processes with Gaussian Processes
arXiv:1606.04753
Abstract
In classical reinforcement learning, when exploring an environment, agents accept arbitrary short term loss for long term gain. This is infeasible for safety critical applications, such as robotics, where even a single unsafe action may cause system failure. In this paper, we address the problem of safely exploring finite Markov decision processes (MDP). We define safety in terms of an, a priori unknown, safety constraint that depends on states and actions. We aim to explore the MDP under this constraint, assuming that the unknown function satisfies regularity conditions expressed via a Gaussian process prior. We develop a novel algorithm for this task and prove that it is able to completely explore the safely reachable part of the MDP without violating the safety constraint. To achieve this, it cautiously explores safe states and actions in order to gain statistical confidence about the safety of unvisited state-action pairs from noisy observations collected while navigating the environment. Moreover, the algorithm explicitly considers reachability when exploring the MDP, ensuring that it does not get stuck in any state with no safe way out. We demonstrate our method on digital terrain models for the task of exploring an unknown map with a rover.
15 pages, extended version with proofs
Cited by in corpus (23)
- Exploration in Deep Reinforcement Learning: A Survey
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- Verifiable Reinforcement Learning via Policy Extraction
- How to Certify Machine Learning Based Safety-critical Systems? A Systematic Literature Review
- Stagewise Safe Bayesian Optimization with Gaussian Processes
- An empirical investigation of the challenges of real-world reinforcement learning
- Safe Reinforcement Learning via Curriculum Induction
- Penalizing side effects using stepwise relative reachability
- Recovery RL: Safe Reinforcement Learning with Learned Recovery Zones
- Cautious Reinforcement Learning with Logical Constraints
- Learning Constraints from Demonstrations
- Transforming task representations to perform novel tasks
- Robust-Adaptive Control of Linear Systems: beyond Quadratic Costs
- Nonstationary Nonparametric Online Learning: Balancing Dynamic Regret and Model Parsimony
- Robust Regression for Safe Exploration in Control
- Provably Correct Training of Neural Network Controllers Using Reachability Analysis
- Online Learning in Kernelized Markov Decision Processes
- Safe Reinforcement Learning with Linear Function Approximation
- Learning to Be Cautious
- Curiosity Killed or Incapacitated the Cat and the Asymptotically Optimal Agent
- Transfer Reinforcement Learning across Homotopy Classes
- Is the Rush to Machine Learning Jeopardizing Safety? Results of a Survey
- Focused Model-Learning and Planning for Non-Gaussian Continuous State-Action Systems