Towards Proving the Adversarial Robustness of Deep Neural Networks
arXiv:1709.02802 · doi:10.4204/EPTCS.257.3
Abstract
Autonomous vehicles are highly complex systems, required to function reliably in a wide variety of situations. Manually crafting software controllers for these vehicles is difficult, but there has been some success in using deep neural networks generated using machine-learning. However, deep neural networks are opaque to human engineers, rendering their correctness very difficult to prove manually; and existing automated techniques, which were not designed to operate on neural networks, fail to scale to large systems. This paper focuses on proving the adversarial robustness of deep neural networks, i.e. proving that small perturbations to a correctly-classified input to the network cannot cause it to be misclassified. We describe some of our recent and ongoing work on verifying the adversarial robustness of networks, and discuss some of the open questions we have encountered and how they might be addressed.
In Proceedings FVAV 2017, arXiv:1709.02126
Cited by in corpus (10)
- Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach
- AI Safety Gridworlds
- Interpreting Neural Networks Using Flip Points
- RobOT: Robustness-Oriented Testing for Deep Learning Systems
- Safe Predictors for Enforcing Input-Output Specifications
- Formal methods and software engineering for DL. Security, safety and productivity for DL systems development
- ROBY: Evaluating the Robustness of a Deep Model by its Decision Boundaries
- Exploiting Vulnerability of Pooling in Convolutional Neural Networks by Strict Layer-Output Manipulation for Adversarial Attacks
- Deterministic Certification to Adversarial Attacks via Bernstein Polynomial Approximation
- Uncovering Discrimination Clusters: Quantifying and Explaining Systematic Fairness Violations