Conservative Safety Critics for Exploration
arXiv:2010.14497
Abstract
Safe exploration presents a major challenge in reinforcement learning (RL): when active data collection requires deploying partially trained policies, we must ensure that these policies avoid catastrophically unsafe regions, while still enabling trial and error learning. In this paper, we target the problem of safe exploration in RL by learning a conservative safety estimate of environment states through a critic, and provably upper bound the likelihood of catastrophic failures at every training iteration. We theoretically characterize the tradeoff between safety and policy improvement, show that the safety constraints are likely to be satisfied with high probability during training, derive provable convergence guarantees for our approach, which is no worse asymptotically than standard RL, and demonstrate the efficacy of the proposed approach on a suite of challenging navigation, manipulation, and locomotion tasks. Empirically, we show that the proposed approach can achieve competitive task performance while incurring significantly lower catastrophic failure rates during training than prior methods. Videos are at this url https://sites.google.com/view/conservative-safety-critics/home
Published as a conference paper in ICLR 2021
References in corpus (17)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Trust Region Policy Optimization
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Conservative Q-Learning for Offline Reinforcement Learning
- Safe Model-based Reinforcement Learning with Stability Guarantees
- Safe Exploration in Continuous Action Spaces
- Reward Constrained Policy Optimization
- robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
- Benchmarking Batch Deep Reinforcement Learning Algorithms
- Lyapunov-based Safe Policy Optimization for Continuous Control
- Constrained Policy Optimization
- On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift
- Responsive Safety in Reinforcement Learning by PID Lagrangian Methods
- Learning to be Safe: Deep RL with a Safety Critic
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
- SAMBA: Safe Model-Based & Active Reinforcement Learning
Cited by in corpus (7)
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- Safe Reinforcement Learning Using Black-Box Reachability Analysis
- TRC: Trust Region Conditional Value at Risk for Safe Reinforcement Learning
- Efficient Off-Policy Safe Reinforcement Learning Using Trust Region Conditional Value at Risk
- Safe Exploration by Solving Early Terminated MDP
- C-Learning: Horizon-Aware Cumulative Accessibility Estimation
- Auditing Robot Learning for Safety and Compliance during Deployment