Attacking the Madry Defense Model with -based Adversarial Examples
arXiv:1710.10733
Abstract
The Madry Lab recently hosted a competition designed to test the robustness of their adversarially trained MNIST model. Attacks were constrained to perturb each pixel of the input image by a scaled maximal distortion = 0.3. This discourages the use of attacks which are not optimized on the distortion metric. Our experimental results demonstrate that by relaxing the constraint of the competition, the elastic-net attack to deep neural networks (EAD) can generate transferable adversarial examples which, despite their high average distortion, have minimal visual distortion. These results call into question the use of as a sole measure for visual distortion, and further demonstrate the power of EAD at generating robust adversarial examples.
Accepted to ICLR 2018 Workshops
References in corpus (3)
Cited by in corpus (45)
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
- On Evaluating Adversarial Robustness
- Natural Adversarial Examples
- Motivating the Rules of the Game for Adversarial Example Research
- Adversarial Examples Are a Natural Consequence of Test Error in Noise
- Adversarial Training and Robustness for Multiple Perturbations
- Adversarial Training against Location-Optimized Adversarial Patches
- Testing Robustness Against Unforeseen Adversaries
- Security and Privacy Issues in Deep Learning
- NATTACK: Learning the Distributions of Adversarial Examples for an Improved Black-Box Attack on Deep Neural Networks
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- An Empirical Evaluation on Robustness and Uncertainty of Regularization Methods
- Confidence-Calibrated Adversarial Training: Generalizing to Unseen Attacks
- Pixle: a fast and effective black-box attack based on rearranging pixels
- Curriculum Adversarial Training
- Bypassing Feature Squeezing by Increasing Adversary Strength
- On the Limitation of Local Intrinsic Dimensionality for Characterizing the Subspaces of Adversarial Examples
- GenAttack: Practical Black-box Attacks with Gradient-Free Optimization
- Transfer of Adversarial Robustness Between Perturbation Types
- Adversarial Training Versus Weight Decay
- Attacking Optical Character Recognition (OCR) Systems with Adversarial Watermarks
- Scratch that! An Evolution-based Adversarial Attack against Neural Networks
- Robustifying Models Against Adversarial Attacks by Langevin Dynamics
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- On the Limitation of MagNet Defense against -based Adversarial Examples
- Distributionally Adversarial Attack
- Using Undervolting as an On-Device Defense Against Adversarial Machine Learning Attacks
- Is Robustness the Cost of Accuracy? -- A Comprehensive Study on the Robustness of 18 Deep Image Classification Models
- The Efficacy of SHIELD under Different Threat Models
- Adversarial Learning Guarantees for Linear Hypotheses and Neural Networks
- CAAD 2018: Generating Transferable Adversarial Examples
- Provable robustness against all adversarial -perturbations for
- Robust and Private Learning of Halfspaces
- Adversarial Defense Through Network Profiling Based Path Extraction
- Towards Deep Learning Models Resistant to Large Perturbations
- Implicit Generative Modeling of Random Noise during Training for Adversarial Robustness
- A cryptographic approach to black box adversarial machine learning
- On The Utility of Conditional Generation Based Mutual Information for Characterizing Adversarial Subspaces
- A Multiclass Boosting Framework for Achieving Fast and Provable Adversarial Robustness
- Adversarial Examples and Metrics
- Adversarial Examples as an Input-Fault Tolerance Problem
- Defense Through Diverse Directions
- Towards Robustness against Unsuspicious Adversarial Examples
- Group-Structured Adversarial Training
- Machine learning pipeline for battery state of health estimation