A Boundary Tilting Persepective on the Phenomenon of Adversarial Examples
arXiv:1608.07690
Abstract
Deep neural networks have been shown to suffer from a surprising weakness: their classification outputs can be changed by small, non-random perturbations of their inputs. This adversarial example phenomenon has been explained as originating from deep networks being "too linear" (Goodfellow et al., 2014). We show here that the linear explanation of adversarial examples presents a number of limitations: the formal argument is not convincing, linear classifiers do not always suffer from the phenomenon, and when they do their adversarial examples are different from the ones affecting deep networks. We propose a new perspective on the phenomenon. We argue that adversarial examples exist when the classification boundary lies close to the submanifold of sampled data, and present a mathematical analysis of this new perspective in the linear case. We define the notion of adversarial strength and show that it can be reduced to the deviation angle between the classifier considered and the nearest centroid classifier. Then, we show that the adversarial strength can be made arbitrarily high independently of the classification performance due to a mechanism that we call boundary tilting. This result leads us to defining a new taxonomy of adversarial examples. Finally, we show that the adversarial strength observed in practice is directly dependent on the level of regularisation used and the strongest adversarial examples, symptomatic of overfitting, can be avoided by using a proper level of regularisation.
arXiv admin note: text overlap with arXiv:1412.6572 by other authors
References in corpus (2)
Cited by in corpus (25)
- A Review of Adversarial Attack and Defense for Classification Methods
- Batch Normalization is a Cause of Adversarial Vulnerability
- Bridging Adversarial Robustness and Gradient Interpretability
- Predify: Augmenting deep neural networks with brain-inspired predictive coding dynamics
- A Simple Fine-tuning Is All You Need: Towards Robust Deep Learning Via Adversarial Fine-tuning
- Towards Understanding Adversarial Examples Systematically: Exploring Data Size, Task and Model Factors
- ML-LOO: Detecting Adversarial Examples with Feature Attribution
- Adversarial Attacks and Defense on Texts: A Survey
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- Sparse Adversarial Video Attacks with Spatial Transformations
- Data from Model: Extracting Data from Non-robust and Robust Models
- Softmax-based Classification is k-means Clustering: Formal Proof, Consequences for Adversarial Attacks, and Improvement through Centroid Based Tailoring
- Understanding Adversarial Examples from the Mutual Influence of Images and Perturbations
- Universal Adversarial Perturbations Through the Lens of Deep Steganography: Towards A Fourier Perspective
- Trust but Verify: An Information-Theoretic Explanation for the Adversarial Fragility of Machine Learning Systems, and a General Defense against Adversarial Attacks
- Adversarial Risk Bounds for Neural Networks through Sparsity based Compression
- What Do Adversarially Robust Models Look At?
- ROBY: Evaluating the Robustness of a Deep Model by its Decision Boundaries
- Who is Responsible for Adversarial Defense?
- An Information-Theoretic Explanation for the Adversarial Fragility of AI Classifiers
- A novel network training approach for open set image recognition
- Analyzing Adversarial Robustness of Deep Neural Networks in Pixel Space: a Semantic Perspective
- Towards A Conceptually Simple Defensive Approach for Few-shot classifiers Against Adversarial Support Samples
- Adversarial Data Encryption
- Hard-label Manifolds: Unexpected Advantages of Query Efficiency for Finding On-manifold Adversarial Examples