Adversarial vulnerability for any classifier
arXiv:1802.08686
Abstract
Despite achieving impressive performance, state-of-the-art classifiers remain highly vulnerable to small, imperceptible, adversarial perturbations. This vulnerability has proven empirically to be very intricate to address. In this paper, we study the phenomenon of adversarial perturbations under the assumption that the data is generated with a smooth generative model. We derive fundamental upper bounds on the robustness to perturbations of any classification function, and prove the existence of adversarial perturbations that transfer well across different classifiers with small risk. Our analysis of the robustness also provides insights onto key properties of generative models, such as their smoothness and dimensionality of latent space. We conclude with numerical experimental results showing that our bounds provide informative baselines to the maximal achievable robustness on several datasets.
NeurIPS 2018
Cited by in corpus (35)
- Adversarial Examples Are Not Bugs, They Are Features
- Robustness May Be at Odds with Accuracy
- Adversarial Examples: Attacks and Defenses for Deep Learning
- A Dual Approach to Scalable Verification of Deep Networks
- Adversarial examples from computational constraints
- Adversarial Training and Robustness for Multiple Perturbations
- Rademacher Complexity for Adversarially Robust Generalization
- Rethinking Softmax Cross-Entropy Loss for Adversarial Robustness
- Adversarially Robust Generalization Just Requires More Unlabeled Data
- A Closer Look at Accuracy vs. Robustness
- Towards Evaluating the Robustness of Deep Diagnostic Models by Adversarial Attack
- Excessive Invariance Causes Adversarial Vulnerability
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- Mixup Inference: Better Exploiting Mixup to Defend Adversarial Attacks
- A Causal View on Robustness of Neural Networks
- Universal Adversarial Audio Perturbations
- Robustness of Bayesian Neural Networks to Gradient-Based Attacks
- The Threat of Adversarial Attacks on Machine Learning in Network Security -- A Survey
- More Data Can Expand the Generalization Gap Between Adversarially Robust and Standard Models
- Randomization matters. How to defend against strong adversarial attacks
- A Spectral View of Adversarially Robust Features
- Adaptive Adversarial Attack on Scene Text Recognition
- Learning Adversarially Robust Representations via Worst-Case Mutual Information Maximization
- Local intrinsic dimensionality estimators based on concentration of measure
- Adversarial Machine Learning Phases of Matter
- Adversarial Risk and Robustness: General Definitions and Implications for the Uniform Distribution
- Robust Large-Margin Learning in Hyperbolic Space
- Robust Adversarial Learning via Sparsifying Front Ends
- Intrinsic Geometric Vulnerability of High-Dimensional Artificial Intelligence
- ATRO: Adversarial Training with a Rejection Option
- Adversarial Robustness Guarantees for Random Deep Neural Networks
- Utilizing Network Properties to Detect Erroneous Inputs
- Empirically Measuring Concentration: Fundamental Limits on Intrinsic Robustness
- Do Deep Minds Think Alike? Selective Adversarial Attacks for Fine-Grained Manipulation of Multiple Deep Neural Networks
- Classifier-independent Lower-Bounds for Adversarial Robustness