Adversarial Examples Detection in Deep Networks with Convolutional Filter Statistics
arXiv:1612.07767
Abstract
Deep learning has greatly improved visual recognition in recent years. However, recent research has shown that there exist many adversarial examples that can negatively impact the performance of such an architecture. This paper focuses on detecting those adversarial examples by analyzing whether they come from the same distribution as the normal examples. Instead of directly training a deep neural network to detect adversarials, a much simpler approach was proposed based on statistics on outputs from convolutional layers. A cascade classifier was designed to efficiently detect adversarials. Furthermore, trained from one particular adversarial generating mechanism, the resulting classifier can successfully detect adversarials from a completely different mechanism as well. The resulting classifier is non-subdifferentiable, hence creates a difficulty for adversaries to attack by using the gradient of the classifier. After detecting adversarial examples, we show that many of them can be recovered by simply performing a small average filter on the image. Those findings should lead to more insights about the classification mechanisms in deep convolutional neural networks.
Published in ICCV 2017
References in corpus (12)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- Going Deeper with Convolutions
- Energy-based Generative Adversarial Network
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Learning Deconvolution Network for Semantic Segmentation
- Stochastic Pooling for Regularization of Deep Convolutional Neural Networks
- Understanding Adversarial Training: Increasing Local Stability of Neural Nets through Robust Optimization
- Learning with a Strong Adversary
- Foveation-based Mechanisms Alleviate Adversarial Examples
- DeepFool: a simple and accurate method to fool deep neural networks
- Learning to Abstain from Binary Prediction
Cited by in corpus (12)
- Characterizing Adversarial Subspaces Using Local Intrinsic Dimensionality
- Defending against Adversarial Images using Basis Functions Transformations
- Detecting Adversarial Attacks on Neural Network Policies with Visual Foresight
- Breaking Transferability of Adversarial Samples with Randomness
- Towards Robust Detection of Adversarial Examples
- Automated Poisoning Attacks and Defenses in Malware Detection Systems: An Adversarial Machine Learning Approach
- ReabsNet: Detecting and Revising Adversarial Examples
- Adversarial Reinforcement Learning under Partial Observability in Autonomous Computer Network Defence
- Reinforcement Learning for Autonomous Defence in Software-Defined Networking
- Clipping free attacks against artificial neural networks
- Detecting Adversarial Samples Using Density Ratio Estimates
- Where Classification Fails, Interpretation Rises