EAD: Elastic-Net Attacks to Deep Neural Networks via Adversarial Examples
arXiv:1709.04114
Abstract
Recent studies have highlighted the vulnerability of deep neural networks (DNNs) to adversarial examples - a visually indistinguishable adversarial image can easily be crafted to cause a well-trained model to misclassify. Existing methods for crafting adversarial examples are based on and distortion metrics. However, despite the fact that distortion accounts for the total variation and encourages sparsity in the perturbation, little has been developed for crafting -based adversarial examples. In this paper, we formulate the process of attacking DNNs via adversarial examples as an elastic-net regularized optimization problem. Our elastic-net attacks to DNNs (EAD) feature -oriented adversarial examples and include the state-of-the-art attack as a special case. Experimental results on MNIST, CIFAR10 and ImageNet show that EAD can yield a distinct set of adversarial examples with small distortion and attains similar attack performance to the state-of-the-art methods in different attack scenarios. More importantly, EAD leads to improved attack transferability and complements adversarial training for DNNs, suggesting novel insights on leveraging distortion in adversarial machine learning and security implications of DNNs.
To be published at AAAI 2018
References in corpus (11)
- Distilling the Knowledge in a Neural Network
- Explaining and Harnessing Adversarial Examples
- Ensemble Adversarial Training: Attacks and Defenses
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks
- ZOO: Zeroth Order Optimization based Black-box Attacks to Deep Neural Networks without Training Substitute Models
- Understanding Black-box Predictions via Influence Functions
- Delving into Transferable Adversarial Examples and Black-box Attacks
- Robust Physical-World Attacks on Deep Learning Models
- Universal adversarial perturbations
- Towards Interpretable Deep Neural Networks by Leveraging Adversarial Examples
- SafetyNet: Detecting and Rejecting Adversarial Examples Robustly
Cited by in corpus (18)
- On Evaluating Adversarial Robustness
- Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach
- Towards Fast Computation of Certified Robustness for ReLU Networks
- Characterizing Audio Adversarial Examples Using Temporal Dependency
- Fault Sneaking Attack: a Stealthy Framework for Misleading Deep Neural Networks
- Detecting Adversarial Samples for Deep Neural Networks through Mutation Testing
- Weighted-Sampling Audio Adversarial Example Attack
- Defend Deep Neural Networks Against Adversarial Examples via Fixed and Dynamic Quantized Activation Functions
- On the Limitation of Local Intrinsic Dimensionality for Characterizing the Subspaces of Adversarial Examples
- What You See is Not What the Network Infers: Detecting Adversarial Examples Based on Semantic Contradiction
- A Panda? No, It's a Sloth: Slowdown Attacks on Adaptive Multi-Exit Neural Network Inference
- Adversarial Detection and Correction by Matching Prediction Distributions
- Appending Adversarial Frames for Universal Video Attack
- Blind Adversarial Network Perturbations
- Robustness Certificates Against Adversarial Examples for ReLU Networks
- Blurring Fools the Network -- Adversarial Attacks by Feature Peak Suppression and Gaussian Blurring
- On The Utility of Conditional Generation Based Mutual Information for Characterizing Adversarial Subspaces
- On Configurable Defense against Adversarial Example Attacks