Rigorous Agent Evaluation: An Adversarial Approach to Uncover Catastrophic Failures
arXiv:1812.01647
Abstract
This paper addresses the problem of evaluating learning systems in safety critical domains such as autonomous driving, where failures can have catastrophic consequences. We focus on two problems: searching for scenarios when learned agents fail and assessing their probability of failure. The standard method for agent evaluation in reinforcement learning, Vanilla Monte Carlo, can miss failures entirely, leading to the deployment of unsafe agents. We demonstrate this is an issue for current agents, where even matching the compute used for training is sometimes insufficient for evaluation. To address this shortcoming, we draw upon the rare event probability estimation literature and propose an adversarial evaluation approach. Our approach focuses evaluation on adversarially chosen situations, while still providing unbiased estimates of failure probabilities. The key difficulty is in identifying these adversarial situations -- since failures are rare there is little signal to drive optimization. To solve this we propose a continuation approach that learns failure modes in related but less robust agents. Our approach also allows reuse of data already collected for training the agent. We demonstrate the efficacy of adversarial evaluation on two standard domains: humanoid control and simulated driving. Experimental results show that our methods can find catastrophic failures and estimate failures rates of agents multiple orders of magnitude faster than standard evaluation schemes, in minutes to hours rather than days.
References in corpus (14)
- Explaining and Harnessing Adversarial Examples
- Towards Deep Learning Models Resistant to Adversarial Attacks
- DeepXplore: Automated Whitebox Testing of Deep Learning Systems
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
- DeepMind Control Suite
- Distributed Prioritized Experience Replay
- On a Formal Model of Safe and Scalable Self-driving Cars
- Distributed Distributional Deterministic Policy Gradients
- Towards the first adversarially robust neural network model on MNIST
- Motivating the Rules of the Game for Adversarial Example Research
- Constructing Unrestricted Adversarial Examples with Generative Models
- Constrained Policy Optimization
- A Lyapunov-based Approach to Safe Reinforcement Learning
- Unrestricted Adversarial Examples
Cited by in corpus (15)
- Scalable agent alignment via reward modeling: a research direction
- A Survey of Algorithms for Black-Box Safety Validation of Cyber-Physical Systems
- How to Certify Machine Learning Based Safety-critical Systems? A Systematic Literature Review
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State Observations
- Analysing Deep Reinforcement Learning Agents Trained with Domain Randomisation
- A Statistical Approach to Assessing Neural Network Robustness
- Predicting Model Failure using Saliency Maps in Autonomous Driving Systems
- Neural Bridge Sampling for Evaluating Safety-Critical Autonomous Systems
- Evaluating the Robustness of Collaborative Agents
- Causal Analysis of Agent Behavior for AI Safety
- Play to Grade: Testing Coding Games as Classifying Markov Decision Process
- Scalable Safety-Critical Policy Evaluation with Accelerated Rare Event Sampling
- Deep Probabilistic Accelerated Evaluation: A Robust Certifiable Rare-Event Simulation Methodology for Black-Box Safety-Critical Systems
- Adversarial Reinforcement Learning in Dynamic Channel Access and Power Control
- CoachNet: An Adversarial Sampling Approach for Reinforcement Learning