Adversarial Spheres
arXiv:1801.02774
Abstract
State of the art computer vision models have been shown to be vulnerable to small adversarial perturbations of the input. In other words, most images in the data distribution are both correctly classified by the model and are very close to a visually similar misclassified image. Despite substantial research interest, the cause of the phenomenon is still poorly understood and remains unsolved. We hypothesize that this counter intuitive behavior is a naturally occurring result of the high dimensional geometry of the data manifold. As a first step towards exploring this hypothesis, we study a simple synthetic dataset of classifying between two concentric high dimensional spheres. For this dataset we show a fundamental tradeoff between the amount of test error and the average distance to nearest error. In particular, we prove that any model which misclassifies a small constant fraction of a sphere will be vulnerable to adversarial perturbations of size . Surprisingly, when we train several different architectures on this dataset, all of their error sets naturally approach this theoretical bound. As a result of the theory, the vulnerability of neural networks to small adversarial perturbations is a logical consequence of the amount of test error observed. We hope that our theoretical analysis of this very simple case will point the way forward to explore how the geometry of complex real-world data sets leads to adversarial examples.
References in corpus (2)
Cited by in corpus (28)
- Adversarial Examples Are Not Bugs, They Are Features
- Robustness May Be at Odds with Accuracy
- Sensitivity and Generalization in Neural Networks: an Empirical Study
- Adversarial examples from computational constraints
- Detecting and Diagnosing Adversarial Images with Class-Conditional Capsule Reconstructions
- Do Wider Neural Networks Really Help Adversarial Robustness?
- Adversarial Robustness via Fisher-Rao Regularization
- More Data Can Expand the Generalization Gap Between Adversarially Robust and Standard Models
- PAC-learning in the presence of evasion adversaries
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- Adversarial Machine Learning Phases of Matter
- Learning Adversarially Robust Representations via Worst-Case Mutual Information Maximization
- Bayesian Adversarial Spheres: Bayesian Inference and Adversarial Examples in a Noiseless Setting
- Adversarial Visual Robustness by Causal Intervention
- Interpreting Adversarial Examples with Attributes
- ColdGANs: Taming Language GANs with Cautious Sampling Strategies
- A Bayes-Optimal View on Adversarial Examples
- Analytical Moment Regularizer for Gaussian Robust Networks
- Connecting Lyapunov Control Theory to Adversarial Attacks
- On Procedural Adversarial Noise Attack And Defense
- Lower Bounds on Cross-Entropy Loss in the Presence of Test-time Adversaries
- Adversarial Examples in Remote Sensing
- Explaining generalization in deep learning: progress and fundamental limits
- Query complexity of adversarial attacks
- Adversarial Data Encryption
- A mathematical theory of imperfect communication: Energy efficiency considerations in multi-level coding
- Interpreting Attributions and Interactions of Adversarial Attacks
- Classifier-independent Lower-Bounds for Adversarial Robustness