Approximate Model-Based Shielding for Safe Reinforcement Learning
arXiv:2308.00707 · doi:10.3233/FAIA230357
Abstract
Reinforcement learning (RL) has shown great potential for solving complex tasks in a variety of domains. However, applying RL to safety-critical systems in the real-world is not easy as many algorithms are sample-inefficient and maximising the standard RL objective comes with no guarantees on worst-case performance. In this paper we propose approximate model-based shielding (AMBS), a principled look-ahead shielding algorithm for verifying the performance of learned RL policies w.r.t. a set of given safety constraints. Our algorithm differs from other shielding approaches in that it does not require prior knowledge of the safety-relevant dynamics of the system. We provide a strong theoretical justification for AMBS and demonstrate superior performance to other safety-aware approaches on a set of Atari games with state-dependent safety-labels.
Accepted at ECAI 2023 (main technical track)
References in corpus (15)
- DeepMind Control Suite
- Dopamine: A Research Framework for Deep Reinforcement Learning
- Benchmarking Batch Deep Reinforcement Learning Algorithms
- Lyapunov-based Safe Policy Optimization for Continuous Control
- Mastering Diverse Domains through World Models
- Shield Synthesis: Runtime Enforcement for Reactive Systems
- IPO: Interior-point Policy Optimization under Constraints
- Learning Belief Representations for Imitation Learning in POMDPs
- Constrained Policy Optimization via Bayesian World Models
- Learning Barrier Certificates: Towards Safe Reinforcement Learning with Zero Training-time Violations
- Safe Reinforcement Learning by Imagining the Near Future
- Do Androids Dream of Electric Fences? Safety-Aware Reinforcement Learning with Latent Shielding
- Learning in POMDPs is Sample-Efficient with Hindsight Observability
- Model-based Dynamic Shielding for Safe and Efficient Multi-Agent Reinforcement Learning
- Approximate Shielding of Atari Agents for Safe Exploration