Adversarial Robustness May Be at Odds With Simplicity
arXiv:1901.00532
Abstract
Current techniques in machine learning are so far are unable to learn classifiers that are robust to adversarial perturbations. However, they are able to learn non-robust classifiers with very high accuracy, even in the presence of random perturbations. Towards explaining this gap, we highlight the hypothesis that In this note, we show that this hypothesis is indeed possible, by giving several theoretical examples of classification tasks and sets of "simple" classifiers for which: (1) There exists a simple classifier with high standard accuracy, and also high accuracy under random noise. (2) Any simple classifier is not robust: it must have high adversarial loss with perturbations. (3) Robust classification is possible, but only with more complex classifiers (exponentially more complex, in some examples). Moreover, This suggests an alternate explanation of this phenomenon, which appears in practice: the tradeoff may occur not because the classification task inherently requires such a tradeoff (as in [Tsipras-Santurkar-Engstrom-Turner-Madry `18]), but because the structure of our current classifiers imposes such a tradeoff.
welcome
Cited by in corpus (17)
- Triple Wins: Boosting Accuracy, Robustness and Efficiency Together by Enabling Input-Adaptive Inference
- Once-for-All Adversarial Training: In-Situ Tradeoff between Robustness and Accuracy for Free
- Precise Tradeoffs in Adversarial Training for Linear Regression
- How benign is benign overfitting?
- Label Smoothing and Adversarial Robustness
- Adversarial Feature Augmentation and Normalization for Visual Recognition
- Computational Limitations in Robust Classification and Win-Win Results
- Adversarial Learning Guarantees for Linear Hypotheses and Neural Networks
- Removing Spurious Features can Hurt Accuracy and Affect Groups Disproportionately
- Robustness, Privacy, and Generalization of Adversarial Training
- A Useful Taxonomy for Adversarial Robustness of Neural Networks
- Classification and Adversarial examples in an Overparameterized Linear Model: A Signal Processing Perspective
- Gödel's Sentence Is An Adversarial Example But Unsolvable
- Unique properties of adversarially trained linear classifiers on Gaussian data
- Estimating Principal Components under Adversarial Perturbations
- Achieving Adversarial Robustness Requires An Active Teacher
- Understanding Adversarial Behavior of DNNs by Disentangling Non-Robust and Robust Components in Performance Metric