Robustness of classifiers to universal perturbations: a geometric perspective
arXiv:1705.09554
Abstract
Deep networks have recently been shown to be vulnerable to universal perturbations: there exist very small image-agnostic perturbations that cause most natural images to be misclassified by such classifiers. In this paper, we propose the first quantitative analysis of the robustness of classifiers to universal perturbations, and draw a formal link between the robustness to universal perturbations, and the geometry of the decision boundary. Specifically, we establish theoretical bounds on the robustness of classifiers under two decision boundary models (flat and curved models). We show in particular that the robustness of deep networks to universal perturbations is driven by a key property of their curvature: there exists shared directions along which the decision boundary of deep networks is systematically positively curved. Under such conditions, we prove the existence of small universal perturbations. Our analysis further provides a novel geometric method for computing universal perturbations, in addition to explaining their properties.
Published at ICLR 2018
Cited by in corpus (29)
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- Arms Race in Adversarial Malware Detection: A Survey
- Universal Adversarial Audio Perturbations
- Universal adversarial examples in speech command classification
- Transferable Universal Adversarial Perturbations Using Generative Models
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- Universal Adversarial Perturbation for Text Classification
- Defending Against Adversarial Attacks by Suppressing the Largest Eigenvalue of Fisher Information Matrix
- Decision-based Universal Adversarial Attack
- Built-in Vulnerabilities to Imperceptible Adversarial Perturbations
- Interpreting Adversarial Examples with Attributes
- Efficient Two-Step Adversarial Defense for Deep Neural Networks
- Why Adversarial Reprogramming Works, When It Fails, and How to Tell the Difference
- On the Effect of Low-Rank Weights on Adversarial Robustness of Neural Networks
- Understanding Adversarial Examples from the Mutual Influence of Images and Perturbations
- Machine vs Machine: Minimax-Optimal Defense Against Adversarial Examples
- Label Universal Targeted Attack
- Butterfly Effect: Bidirectional Control of Classification Performance by Small Additive Perturbation
- Global Adversarial Attacks for Assessing Deep Learning Robustness
- Adversarial Risk Bounds for Neural Networks through Sparsity based Compression
- A Game Theoretic Analysis of Additive Adversarial Attacks and Defenses
- Deep neural network loses attention to adversarial images
- Heating up decision boundaries: isocapacitory saturation, adversarial scenarios and generalization bounds
- On the Structural Sensitivity of Deep Convolutional Networks to the Directions of Fourier Basis Functions
- On Universalized Adversarial and Invariant Perturbations
- Adaptive Gradient for Adversarial Perturbations Generation
- An Explainable Adversarial Robustness Metric for Deep Learning Neural Networks
- Analyzing Adversarial Robustness of Deep Neural Networks in Pixel Space: a Semantic Perspective
- Defense-guided Transferable Adversarial Attacks