Fast Feature Fool: A data independent approach to universal adversarial perturbations
arXiv:1707.05572
Abstract
State-of-the-art object recognition Convolutional Neural Networks (CNNs) are shown to be fooled by image agnostic perturbations, called universal adversarial perturbations. It is also observed that these perturbations generalize across multiple networks trained on the same target data. However, these algorithms require training data on which the CNNs were trained and compute adversarial perturbations via complex optimization. The fooling performance of these approaches is directly proportional to the amount of available training data. This makes them unsuitable for practical attacks since its unreasonable for an attacker to have access to the training data. In this paper, for the first time, we propose a novel data independent approach to generate image agnostic perturbations for a range of CNNs trained for object recognition. We further show that these perturbations are transferable across multiple network architectures trained either on same or different data. In the absence of data, our method generates universal adversarial perturbations efficiently via fooling the features learned at multiple layers thereby causing CNNs to misclassify. Experiments demonstrate impressive fooling rates and surprising transferability for the proposed universal perturbations generated without any training data.
BMVC 2017 and codes are available at https://github.com/utsavgarg/fast-feature-fool
References in corpus (5)
Cited by in corpus (18)
- Exploring the Space of Black-box Attacks on Deep Neural Networks
- Adversarial Examples in Modern Machine Learning: A Review
- The Vulnerability of Semantic Segmentation Networks to Adversarial Attacks in Autonomous Driving: Enhancing Extensive Environment Sensing
- Universal Adversarial Perturbations: A Survey
- T-Miner: A Generative Approach to Defend Against Trojan Attacks on DNN-based Text Classification
- Universal Adversarial Perturbation for Text Classification
- Convergence and Margin of Adversarial Training on Separable Data
- Who's Afraid of Adversarial Queries? The Impact of Image Modifications on Content-based Image Retrieval
- Performance Evaluation of Adversarial Attacks: Discrepancies and Solutions
- A Method for Computing Class-wise Universal Adversarial Perturbations
- Understanding Adversarial Examples from the Mutual Influence of Images and Perturbations
- Label Universal Targeted Attack
- A Self-supervised Approach for Adversarial Robustness
- Butterfly Effect: Bidirectional Control of Classification Performance by Small Additive Perturbation
- Real-time Detection of Practical Universal Adversarial Perturbations
- Deep Bayesian Image Set Classification: A Defence Approach against Adversarial Attacks
- Adversarial Attacks with Time-Scale Representations
- Attack to Fool and Explain Deep Networks