Universal adversarial perturbations
arXiv:1610.08401
Abstract
Given a state-of-the-art deep neural network classifier, we show the existence of a universal (image-agnostic) and very small perturbation vector that causes natural images to be misclassified with high probability. We propose a systematic algorithm for computing universal perturbations, and show that state-of-the-art deep neural networks are highly vulnerable to such perturbations, albeit being quasi-imperceptible to the human eye. We further empirically analyze these universal perturbations and show, in particular, that they generalize very well across neural networks. The surprising existence of universal perturbations reveals important geometric correlations among the high-dimensional decision boundary of classifiers. It further outlines potential security breaches with the existence of single directions in the input space that adversaries can possibly exploit to break a classifier on most natural images.
Accepted at IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
References in corpus (5)
Cited by in corpus (30)
- Adversarial Examples, Uncertainty, and Transfer Testing Robustness in Gaussian Process Hybrid Deep Networks
- Fast Feature Fool: A data independent approach to universal adversarial perturbations
- Adversarial Examples that Fool Detectors
- UPSET and ANGRI : Breaking High Performance Image Classifiers
- Blocking Transferability of Adversarial Examples in Black-Box Learning Systems
- Exploring the Space of Black-box Attacks on Deep Neural Networks
- A Survey of Game Theoretic Approaches for Adversarial Machine Learning in Cybersecurity Tasks
- Contrastive Learning with Adversarial Examples
- Detecting Adversarial Examples by Input Transformations, Defense Perturbations, and Voting
- Standard detectors aren't (currently) fooled by physical adversarial stop signs
- Adversarially Robust Neural Architectures
- Confidence estimation in Deep Neural networks via density modelling
- On the Convergence of SGD with Biased Gradients
- Art of singular vectors and universal adversarial perturbations
- LatentPoison - Adversarial Attacks On The Latent Space
- SafeAMC: Adversarial training for robust modulation recognition models
- Patch-wise++ Perturbation for Adversarial Targeted Attacks
- Robustness Certificates Against Adversarial Examples for ReLU Networks
- Verification of Binarized Neural Networks via Inter-Neuron Factoring
- HeNet: A Deep Learning Approach on Intel Processor Trace for Effective Exploit Detection
- A Singular Value Perspective on Model Robustness
- Measurement-driven Security Analysis of Imperceptible Impersonation Attacks
- Butterfly Effect: Bidirectional Control of Classification Performance by Small Additive Perturbation
- Adequacy of the Gradient-Descent Method for Classifier Evasion Attacks
- Benchmarking Popular Classification Models' Robustness to Random and Targeted Corruptions
- Strategies for Robust Image Classification
- Where Classification Fails, Interpretation Rises
- A game-theoretic analysis of DoS attacks on driverless vehicles
- P-CapsNets: a General Form of Convolutional Neural Networks
- Adversarial Attacks with Time-Scale Representations