Provably Minimally-Distorted Adversarial Examples
arXiv:1709.10207
Abstract
The ability to deploy neural networks in real-world, safety-critical systems is severely limited by the presence of adversarial examples: slightly perturbed inputs that are misclassified by the network. In recent years, several techniques have been proposed for increasing robustness to adversarial examples --- and yet most of these have been quickly shown to be vulnerable to future attacks. For example, over half of the defenses proposed by papers accepted at ICLR 2018 have already been broken. We propose to address this difficulty through formal verification techniques. We show how to construct provably minimally distorted adversarial examples: given an arbitrary neural network and input sample, we can construct adversarial examples which we prove are of minimal distortion. Using this approach, we demonstrate that one of the recent ICLR defense proposals, adversarial retraining, provably succeeds at increasing the distortion required to construct adversarial examples by a factor of 4.2.
References in corpus (4)
Cited by in corpus (44)
- Adversarial Risk and the Dangers of Evaluating Against Weak Attacks
- Efficient Neural Network Robustness Certification with General Activation Functions
- Formal Security Analysis of Neural Networks using Symbolic Intervals
- On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models
- Adversarial Examples: Attacks and Defenses for Deep Learning
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- A Game-Based Approximate Verification of Deep Neural Networks with Provable Guarantees
- Security and Privacy Issues in Deep Learning
- Tight Certificates of Adversarial Robustness for Randomly Smoothed Classifiers
- Understanding and Improving Fast Adversarial Training
- Certified Robustness for Top-k Predictions against Adversarial Perturbations via Randomized Smoothing
- L2-Nonexpansive Neural Networks
- Harnessing the Vulnerability of Latent Layers in Adversarially Trained Models
- Global Robustness Evaluation of Deep Neural Networks with Provable Guarantees for the Norm
- Black-Box Certification with Randomized Smoothing: A Functional Optimization Based Framework
- Evading classifiers in discrete domains with provable optimality guarantees
- Achieving Verified Robustness to Symbol Substitutions via Interval Bound Propagation
- Biologically Inspired Mechanisms for Adversarial Robustness
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- Second-Order Provable Defenses against Adversarial Attacks
- Identify Susceptible Locations in Medical Records via Adversarial Attacks on Deep Predictive Models
- PointCloud Saliency Maps
- Improving the Tightness of Convex Relaxation Bounds for Training Certifiably Robust Classifiers
- ART: Abstraction Refinement-Guided Training for Provably Correct Neural Networks
- Built-in Vulnerabilities to Imperceptible Adversarial Perturbations
- Almost Tight L0-norm Certified Robustness of Top-k Predictions against Adversarial Perturbations
- A Method for Computing Class-wise Universal Adversarial Perturbations
- Certified Robustness of Graph Neural Networks against Adversarial Structural Perturbation
- A randomized gradient-free attack on ReLU networks
- Tight Second-Order Certificates for Randomized Smoothing
- A Survey on Trust Metrics for Autonomous Robotic Systems
- Formal methods and software engineering for DL. Security, safety and productivity for DL systems development
- Active Learning Under Malicious Mislabeling and Poisoning Attacks
- A Game Theoretic Analysis of Additive Adversarial Attacks and Defenses
- IWA: Integrated Gradient based White-box Attacks for Fooling Deep Neural Networks
- Insta-RS: Instance-wise Randomized Smoothing for Improved Robustness and Accuracy
- Fast and Stable Interval Bounds Propagation for Training Verifiably Robust Models
- Incorrect by Construction: Fine Tuning Neural Networks for Guaranteed Performance on Finite Sets of Examples
- A New Angle on L2 Regularization
- A Useful Taxonomy for Adversarial Robustness of Neural Networks
- Output Randomization: A Novel Defense for both White-box and Black-box Adversarial Models
- Learning to Separate Clusters of Adversarial Representations for Robust Adversarial Detection
- A Primer on Multi-Neuron Relaxation-based Adversarial Robustness Certification
- Thwarting finite difference adversarial attacks with output randomization