Adversarial Examples Are a Natural Consequence of Test Error in Noise
arXiv:1901.10513
Abstract
Over the last few years, the phenomenon of adversarial examples --- maliciously constructed inputs that fool trained machine learning models --- has captured the attention of the research community, especially when the adversary is restricted to small modifications of a correctly handled input. Less surprisingly, image classifiers also lack human-level performance on randomly corrupted images, such as images with additive Gaussian noise. In this paper we provide both empirical and theoretical evidence that these are two manifestations of the same underlying phenomenon, establishing close connections between the adversarial robustness and corruption robustness research programs. This suggests that improving adversarial robustness should go hand in hand with improving performance in the presence of more general and realistic image corruptions. Based on our results we recommend that future adversarial defenses consider evaluating the robustness of their methods to distributional shift with benchmarks such as Imagenet-C.
References in corpus (2)
Cited by in corpus (32)
- On Evaluating Adversarial Robustness
- Improving Robustness Without Sacrificing Accuracy with Patch Gaussian Augmentation
- Conditional GAN for timeseries generation
- Batch Normalization is a Cause of Adversarial Vulnerability
- Test-Time Adaptation to Distribution Shift by Confidence Maximization and Input Transformation
- BREEDS: Benchmarks for Subpopulation Shift
- Towards an Adversarially Robust Normalization Approach
- Towards Understanding Adversarial Examples Systematically: Exploring Data Size, Task and Model Factors
- Investigating Vulnerability to Adversarial Examples on Multimodal Data Fusion in Deep Learning
- 3DB: A Framework for Debugging Computer Vision Models
- Using learned optimizers to make models robust to input noise
- Convergence and Margin of Adversarial Training on Separable Data
- Lower Bounds for Adversarially Robust PAC Learning
- Jacobian Adversarially Regularized Networks for Robustness
- Fighting Gradients with Gradients: Dynamic Defenses against Adversarial Attacks
- Reliable and Trustworthy Machine Learning for Health Using Dataset Shift Detection
- Stateful Detection of Black-Box Adversarial Attacks
- Understanding Adversarial Examples from the Mutual Influence of Images and Perturbations
- Training on Test Data with Bayesian Adaptation for Covariate Shift
- Predicting with High Correlation Features
- Adversarial Heart Attack: Neural Networks Fooled to Segment Heart Symbols in Chest X-Ray Images
- Adversarial and Natural Perturbations for General Robustness
- Natural Perturbed Training for General Robustness of Neural Network Classifiers
- Adversarial robustness via stochastic regularization of neural activation sensitivity
- Robusta: Robust AutoML for Feature Selection via Reinforcement Learning
- Classification and Adversarial examples in an Overparameterized Linear Model: A Signal Processing Perspective
- Contextual Fusion For Adversarial Robustness
- Reject Illegal Inputs with Generative Classifier Derived from Any Discriminative Classifier
- Countering Adversarial Examples: Combining Input Transformation and Noisy Training
- Robustifying Adversarial Training to the Union of Perturbation Models
- Adversarial Data Encryption
- Gödel's Sentence Is An Adversarial Example But Unsolvable