DeepFool: a simple and accurate method to fool deep neural networks
arXiv:1511.04599
Abstract
State-of-the-art deep neural networks have achieved impressive results on many image classification tasks. However, these same architectures have been shown to be unstable to small, well sought, perturbations of the images. Despite the importance of this phenomenon, no effective methods have been proposed to accurately compute the robustness of state-of-the-art deep classifiers to such perturbations on large-scale datasets. In this paper, we fill this gap and propose the DeepFool algorithm to efficiently compute perturbations that fool deep networks, and thus reliably quantify the robustness of these classifiers. Extensive experimental results show that our approach outperforms recent methods in the task of computing adversarial perturbations and making classifiers more robust.
In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016
References in corpus (2)
Cited by in corpus (57)
- Robust Physical-World Attacks on Deep Learning Models
- Generating Adversarial Examples with Adversarial Networks
- MagNet: a Two-Pronged Defense against Adversarial Examples
- Houdini: Fooling Deep Structured Prediction Models
- Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models
- Adversarial Examples that Fool Detectors
- Physical Adversarial Examples for Object Detectors
- Provably Minimally-Distorted Adversarial Examples
- Universal adversarial perturbations
- Towards Interpretable Deep Neural Networks by Leveraging Adversarial Examples
- Towards Verified Artificial Intelligence
- APE-GAN: Adversarial Perturbation Elimination with GAN
- Testing Robustness Against Unforeseen Adversaries
- Exploring the Space of Black-box Attacks on Deep Neural Networks
- Contrastive Learning with Adversarial Examples
- Adversarial Examples in Modern Machine Learning: A Review
- Standard detectors aren't (currently) fooled by physical adversarial stop signs
- Detecting Adversarial Examples by Input Transformations, Defense Perturbations, and Voting
- SafetyNet: Detecting and Rejecting Adversarial Examples Robustly
- Adversarially Robust Neural Architectures
- Quality Resilient Deep Neural Networks
- Rogue Signs: Deceiving Traffic Sign Recognition with Malicious Ads and Logos
- A Survey of Black-Box Adversarial Attacks on Computer Vision Models
- Attacking Visual Language Grounding with Adversarial Examples: A Case Study on Neural Image Captioning
- White Paper Machine Learning in Certified Systems
- PeerNets: Exploiting Peer Wisdom Against Adversarial Attacks
- Defending Against Adversarial Examples with K-Nearest Neighbor
- Robustness of Rotation-Equivariant Networks to Adversarial Perturbations
- ConvNets and ImageNet Beyond Accuracy: Understanding Mistakes and Uncovering Biases
- Adversarial Examples Detection in Deep Networks with Convolutional Filter Statistics
- Causal Interpretability for Machine Learning -- Problems, Methods and Evaluation
- Adversarial Attacks on Convolutional Neural Networks in Facial Recognition Domain
- The gap between theory and practice in function approximation with deep neural networks
- ATHENA: A Framework based on Diverse Weak Defenses for Building Adversarial Defense
- Fine-grained Synthesis of Unrestricted Adversarial Examples
- Geometric robustness of deep networks: analysis and improvement
- A Simple Cache Model for Image Recognition
- Testing Deep Learning Models for Image Analysis Using Object-Relevant Metamorphic Relations
- Differentiable Language Model Adversarial Attacks on Categorical Sequence Classifiers
- One Bit Matters: Understanding Adversarial Examples as the Abuse of Redundancy
- Generating Adversarial Examples with Graph Neural Networks
- A Singular Value Perspective on Model Robustness
- Adversarial Defense Through Network Profiling Based Path Extraction
- Butterfly Effect: Bidirectional Control of Classification Performance by Small Additive Perturbation
- Utilizing Network Properties to Detect Erroneous Inputs
- Regularized Ensembles and Transferability in Adversarial Learning
- A principled approach for generating adversarial images under non-smooth dissimilarity metrics
- DeepConsensus: using the consensus of features from multiple layers to attain robust image classification
- A Hierarchical Feature Constraint to Camouflage Medical Adversarial Attacks
- Non-Determinism in Neural Networks for Adversarial Robustness
- Affine Disentangled GAN for Interpretable and Robust AV Perception
- Improving Adversarial Robustness for Free with Snapshot Ensemble
- FUNN: Flexible Unsupervised Neural Network
- Boosting the Robustness Verification of DNN by Identifying the Achilles's Heel
- Generative Adversarial Networks for Black-Box API Attacks with Limited Training Data
- HAWKEYE: Adversarial Example Detector for Deep Neural Networks
- On Configurable Defense against Adversarial Example Attacks