AI Safety Gridworlds
arXiv:1711.09883
Abstract
We present a suite of reinforcement learning environments illustrating various safety properties of intelligent agents. These problems include safe interruptibility, avoiding side effects, absent supervisor, reward gaming, safe exploration, as well as robustness to self-modification, distributional shift, and adversaries. To measure compliance with the intended safe behavior, we equip each environment with a performance function that is hidden from the agent. This allows us to categorize AI safety problems into robustness and specification problems, depending on whether the performance function corresponds to the observed reward function. We evaluate A2C and Rainbow, two recent deep reinforcement learning agents, on our environments and show that they are not able to solve them satisfactorily.
References in corpus (10)
- Explaining and Harnessing Adversarial Examples
- Towards A Rigorous Science of Interpretable Machine Learning
- A Berkeley View of Systems Challenges for AI
- Towards Proving the Adversarial Robustness of Deep Neural Networks
- Trial without Error: Towards Safe Reinforcement Learning via Human Intervention
- Constrained Policy Optimization
- Safe Policy Improvement by Minimizing Robust Baseline Regret
- Low Impact Artificial Intelligences
- APRIL: Active Preference-learning based Reinforcement Learning
- Sequential Extensions of Causal and Evidential Decision Theory
Cited by in corpus (12)
- The Role of Cooperation in Responsible AI Development
- Safety-Guided Deep Reinforcement Learning via Online Gaussian Process Estimation
- A Perspective on Objects and Systematic Generalization in Model-Based RL
- Improving Safety in Reinforcement Learning Using Model-Based Architectures and Human Intervention
- Generalizing from a few environments in safety-critical reinforcement learning
- A Human-Centered Approach to Interactive Machine Learning
- Safer Deep RL with Shallow MCTS: A Case Study in Pommerman
- Parenting: Safe Reinforcement Learning from Human Input
- Categorizing Wireheading in Partially Embedded Agents
- Don't Forget Your Teacher: A Corrective Reinforcement Learning Framework
- Defining Admissible Rewards for High Confidence Policy Evaluation
- Learning the Arrow of Time