The Limitations of Deep Learning in Adversarial Settings
arXiv:1511.07528
Abstract
Deep learning takes advantage of large datasets and computationally efficient training algorithms to outperform other approaches at various machine learning tasks. However, imperfections in the training phase of deep neural networks make them vulnerable to adversarial samples: inputs crafted by adversaries with the intent of causing deep neural networks to misclassify. In this work, we formalize the space of adversaries against deep neural networks (DNNs) and introduce a novel class of algorithms to craft adversarial samples based on a precise understanding of the mapping between inputs and outputs of DNNs. In an application to computer vision, we show that our algorithms can reliably produce samples correctly classified by human subjects but misclassified in specific targets by a DNN with a 97% adversarial success rate while only modifying on average 4.02% of the input features per sample. We then evaluate the vulnerability of different sample classes to adversarial perturbations by defining a hardness measure. Finally, we describe preliminary work outlining defenses against adversarial samples by defining a predictive measure of distance between a benign input and a target classification.
Accepted to the 1st IEEE European Symposium on Security & Privacy, IEEE 2016. Saarbrucken, Germany
References in corpus (3)
Cited by in corpus (39)
- Generating Natural Adversarial Examples
- Adversarial Frontier Stitching for Remote Neural Network Watermarking
- AdvHat: Real-world adversarial attack on ArcFace Face ID system
- Februus: Input Purification Defense Against Trojan Attacks on Deep Neural Network Systems
- Keeping the Bad Guys Out: Protecting and Vaccinating Deep Learning with JPEG Compression
- Defense Against Adversarial Attacks Using Feature Scattering-based Adversarial Training
- Deceiving End-to-End Deep Learning Malware Detectors using Adversarial Examples
- Physical Adversarial Examples for Object Detectors
- Biologically inspired protection of deep networks from adversarial attacks
- Analyzing the Robustness of Nearest Neighbors to Adversarial Examples
- APE-GAN: Adversarial Perturbation Elimination with GAN
- Invisible Mask: Practical Attacks on Face Recognition with Infrared
- Training individually fair ML models with Sensitive Subspace Robustness
- Adversarial Examples in Modern Machine Learning: A Review
- Machine Learning (In) Security: A Stream of Problems
- Reachability Analysis of Deep Neural Networks with Provable Guarantees
- Understanding the One-Pixel Attack: Propagation Maps and Locality Analysis
- Investigating the significance of adversarial attacks and their relation to interpretability for radar-based human activity recognition systems
- The RFML Ecosystem: A Look at the Unique Challenges of Applying Deep Learning to Radio Frequency Applications
- Attack and Defense of Dynamic Analysis-Based, Adversarial Neural Malware Classification Models
- RED-Attack: Resource Efficient Decision based Attack for Machine Learning
- Causal Interpretability for Machine Learning -- Problems, Methods and Evaluation
- Towards Robust Detection of Adversarial Examples
- Multitask Learning Strengthens Adversarial Robustness
- Universal Adversarial Perturbation for Text Classification
- Perturbation Analysis of Gradient-based Adversarial Attacks
- Blind Pre-Processing: A Robust Defense Method Against Adversarial Examples
- Adversarial Attacks in Sound Event Classification
- Adversarial Open-World Person Re-Identification
- Detecting Adversarial Perturbations with Saliency
- Designing Adversarially Resilient Classifiers using Resilient Feature Engineering
- Utilizing Network Properties to Detect Erroneous Inputs
- Rethinking Empirical Evaluation of Adversarial Robustness Using First-Order Attack Methods
- Towards Speeding up Adversarial Training in Latent Spaces
- Adversarial Perturbation Intensity Achieving Chosen Intra-Technique Transferability Level for Logistic Regression
- Can Intelligent Hyperparameter Selection Improve Resistance to Adversarial Examples?
- Game Theory for Adversarial Attacks and Defenses
- SINVAD: Search-based Image Space Navigation for DNN Image Classifier Test Input Generation
- On Configurable Defense against Adversarial Example Attacks