Adversarial Example Detection for DNN Models: A Review and Experimental Comparison
arXiv:2105.00203 · doi:10.1007/s10462-021-10125-w
Abstract
Deep learning (DL) has shown great success in many human-related tasks, which has led to its adoption in many computer vision based applications, such as security surveillance systems, autonomous vehicles and healthcare. Such safety-critical applications have to draw their path to success deployment once they have the capability to overcome safety-critical challenges. Among these challenges are the defense against or/and the detection of the adversarial examples (AEs). Adversaries can carefully craft small, often imperceptible, noise called perturbations to be added to the clean image to generate the AE. The aim of AE is to fool the DL model which makes it a potential risk for DL applications. Many test-time evasion attacks and countermeasures,i.e., defense or detection methods, are proposed in the literature. Moreover, few reviews and surveys were published and theoretically showed the taxonomy of the threats and the countermeasure methods with little focus in AE detection methods. In this paper, we focus on image classification task and attempt to provide a survey for detection methods of test-time evasion attacks on neural network classifiers. A detailed discussion for such methods is provided with experimental results for eight state-of-the-art detectors under different scenarios on four datasets. We also provide potential challenges and future perspectives for this research direction.
Accepted and published in Artificial Intelligence Review journal
References in corpus (21)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Explaining and Harnessing Adversarial Examples
- Fully Convolutional Networks for Semantic Segmentation
- Striving for Simplicity: The All Convolutional Net
- ZOO: Zeroth Order Optimization based Black-box Attacks to Deep Neural Networks without Training Substitute Models
- Evasion Attacks against Machine Learning at Test Time
- Delving into Transferable Adversarial Examples and Black-box Attacks
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Adversarial Attacks on Deep Learning Models in Natural Language Processing: A Survey
- Defensive Distillation is Not Robust to Adversarial Examples
- Adversarial Transformation Networks: Learning to Generate Adversarial Examples
- On Detecting Adversarial Perturbations
- NO Need to Worry about Adversarial Examples in Object Detection in Autonomous Vehicles
- MagNet and "Efficient Defenses Against Adversarial Attacks" are Not Robust to Adversarial Examples
- Early Methods for Detecting Adversarial Images
- Biologically inspired protection of deep networks from adversarial attacks
- UPSET and ANGRI : Breaking High Performance Image Classifiers
- Blocking Transferability of Adversarial Examples in Black-Box Learning Systems
- A Survey of Game Theoretic Approaches for Adversarial Machine Learning in Cybersecurity Tasks
- RAID: Randomized Adversarial-Input Detection for Neural Networks
- Detecting Adversarial Examples and Other Misclassifications in Neural Networks by Introspection
Cited by in corpus (7)
- Survey on Adversarial Attack and Defense for Medical Image Analysis: Methods and Challenges
- Adversarial Attacks and Defenses in Fault Detection and Diagnosis: A Comprehensive Benchmark on the Tennessee Eastman Process
- Topological safeguard for evasion attack interpreting the neural networks' behavior
- Adversarial Artifact Detection in EEG-Based Brain-Computer Interfaces
- How stealthy is stealthy? Studying the Efficacy of Black-Box Adversarial Attacks in the Real World
- Superpixel Attack: Enhancing Black-box Adversarial Attack with Image-driven Division Areas
- Adversarial Detection by Approximation of Ensemble Boundary