Defense against Universal Adversarial Perturbations
arXiv:1711.05929
Abstract
Recent advances in Deep Learning show the existence of image-agnostic quasi-imperceptible perturbations that when applied to `any' image can fool a state-of-the-art network classifier to change its prediction about the image label. These `Universal Adversarial Perturbations' pose a serious threat to the success of Deep Learning in practice. We present the first dedicated framework to effectively defend the networks against such perturbations. Our approach learns a Perturbation Rectifying Network (PRN) as `pre-input' layers to a targeted model, such that the targeted model needs no modification. The PRN is learned from real and synthetic image-agnostic perturbations, where an efficient method to compute the latter is also proposed. A perturbation detector is separately trained on the Discrete Cosine Transform of the input-output difference of the PRN. A query image is first passed through the PRN and verified by the detector. If a perturbation is detected, the output of the PRN is used for label prediction instead of the actual image. A rigorous evaluation shows that our framework can defend the network classifiers against unseen adversarial perturbations in the real-world scenarios with up to 97.5% success rate. The PRN also generalizes well in the sense that training for one targeted network defends another network with a comparable success rate.
Accepted in IEEE CVPR 2018
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Delving into Transferable Adversarial Examples and Black-box Attacks
- A study of the effect of JPG compression on adversarial images
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- A Boundary Tilting Persepective on the Phenomenon of Adversarial Examples
- Adversarial Examples for Semantic Image Segmentation
Cited by in corpus (9)
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- Defending against Adversarial Images using Basis Functions Transformations
- A Noise-Sensitivity-Analysis-Based Test Prioritization Technique for Deep Neural Networks
- Gradient Band-based Adversarial Training for Generalized Attack Immunity of A3C Path Finding
- Towards Security Threats of Deep Learning Systems: A Survey
- ATMPA: Attacking Machine Learning-based Malware Visualization Detection Methods via Adversarial Examples
- Defense-VAE: A Fast and Accurate Defense against Adversarial Attacks
- FineFool: Fine Object Contour Attack via Attention
- Dependency Decomposition and a Reject Option for Explainable Models