On Evaluating Adversarial Robustness
arXiv:1902.06705
Abstract
Correctly evaluating defenses against adversarial examples has proven to be extremely difficult. Despite the significant amount of recent work attempting to design defenses that withstand adaptive attacks, few have succeeded; most papers that propose defenses are quickly shown to be incorrect. We believe a large contributing factor is the difficulty of performing security evaluations. In this paper, we discuss the methodological foundations, review commonly accepted best practices, and suggest new methods for evaluating defenses to adversarial examples. We hope that both researchers developing defenses as well as readers and reviewers who wish to understand the completeness of an evaluation consider our advice in order to avoid common pitfalls.
Living document; source available at https://github.com/evaluating-adversarial-robustness/adv-eval-paper/
References in corpus (8)
- Explaining and Harnessing Adversarial Examples
- ZOO: Zeroth Order Optimization based Black-box Attacks to Deep Neural Networks without Training Substitute Models
- Delving into Transferable Adversarial Examples and Black-box Attacks
- Defensive Distillation is Not Robust to Adversarial Examples
- MagNet and "Efficient Defenses Against Adversarial Attacks" are Not Robust to Adversarial Examples
- Adversarial Examples Are a Natural Consequence of Test Error in Noise
- Adversarial Example Defenses: Ensembles of Weak Defenses are not Strong
- Is AmI (Attacks Meet Interpretability) Robust to Adversarial Examples?
Cited by in corpus (86)
- Fast is better than free: Revisiting adversarial training
- Towards Evaluating the Robustness of Deep Diagnostic Models by Adversarial Attack
- An Alternative Surrogate Loss for PGD-based Adversarial Testing
- Adversarial Examples in Modern Machine Learning: A Review
- Theoretical evidence for adversarial robustness through randomization
- Transfer of Adversarial Robustness Between Perturbation Types
- The RFML Ecosystem: A Look at the Unique Challenges of Applying Deep Learning to Radio Frequency Applications
- Towards Understanding Fast Adversarial Training
- Benchmarking Adversarial Robustness
- Exploiting Excessive Invariance caused by Norm-Bounded Adversarial Robustness
- On Fast Adversarial Robustness Adaptation in Model-Agnostic Meta-Learning
- FenceBox: A Platform for Defeating Adversarial Examples with Data Augmentation Techniques
- Deflecting Adversarial Attacks
- Biologically Inspired Mechanisms for Adversarial Robustness
- Mitigating Advanced Adversarial Attacks with More Advanced Gradient Obfuscation Techniques
- RAID: Randomized Adversarial-Input Detection for Neural Networks
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- Semantic Equivalent Adversarial Data Augmentation for Visual Question Answering
- Evaluating the Robustness of Bayesian Neural Networks Against Different Types of Attacks
- Stateful Detection of Model Extraction Attacks
- Certification of embedded systems based on Machine Learning: A survey
- WaveGuard: Understanding and Mitigating Audio Adversarial Examples
- Gradient Masking and the Underestimated Robustness Threats of Differential Privacy in Deep Learning
- Jacobian Adversarially Regularized Networks for Robustness
- Adversarial Visual Robustness by Causal Intervention
- Now You See It, Now You Dont: Adversarial Vulnerabilities in Computational Pathology
- Softmax-based Classification is k-means Clustering: Formal Proof, Consequences for Adversarial Attacks, and Improvement through Centroid Based Tailoring
- A critique of the DeepSec Platform for Security Analysis of Deep Learning Models
- Adversarial Security Attacks and Perturbations on Machine Learning and Deep Learning Methods
- Interpreting Adversarial Examples with Attributes
- Techniques for Symbol Grounding with SATNet
- Stateful Detection of Black-Box Adversarial Attacks
- Local Reweighting for Adversarial Training
- Towards Improving Adversarial Training of NLP Models
- On Adversarial Robustness of 3D Point Cloud Classification under Adaptive Attacks
- Regularizers for Single-step Adversarial Training
- MagDR: Mask-guided Detection and Reconstruction for Defending Deepfakes
- RobOT: Robustness-Oriented Testing for Deep Learning Systems
- Input Hessian Regularization of Neural Networks
- Adversarially Optimized Mixup for Robust Classification
- Auditing AI models for Verified Deployment under Semantic Specifications
- Benchmarking adversarial attacks and defenses for time-series data
- Ptolemy: Architecture Support for Robust Deep Learning
- Bandlimiting Neural Networks Against Adversarial Attacks
- LAFEAT: Piercing Through Adversarial Defenses with Latent Features
- A unified view on differential privacy and robustness to adversarial examples
- Perception Improvement for Free: Exploring Imperceptible Black-box Adversarial Attacks on Image Classification
- Audit and Assurance of AI Algorithms: A framework to ensure ethical algorithmic practices in Artificial Intelligence
- Adaptive Feature Alignment for Adversarial Training
- Understanding the Error in Evaluating Adversarial Robustness
- On the Exploitability of Audio Machine Learning Pipelines to Surreptitious Adversarial Examples
- On Procedural Adversarial Noise Attack And Defense
- Moving Target Defense for Deep Visual Sensing against Adversarial Examples
- Output Randomization: A Novel Defense for both White-box and Black-box Adversarial Models
- Adversarial robustness via stochastic regularization of neural activation sensitivity
- Who is Responsible for Adversarial Defense?
- Bio-inspired Robustness: A Review
- Usable Security for ML Systems in Mental Health: A Framework
- Building Compact and Robust Deep Neural Networks with Toeplitz Matrices
- An Empirical Study of DNNs Robustification Inefficacy in Protecting Visual Recommenders
- Sampling Prediction-Matching Examples in Neural Networks: A Probabilistic Programming Approach
- Utilizing Adversarial Targeted Attacks to Boost Adversarial Robustness
- The Effects of Image Distribution and Task on Adversarial Robustness
- "What's in the box?!": Deflecting Adversarial Attacks by Randomly Deploying Adversarially-Disjoint Models
- Energy Attack: On Transferring Adversarial Examples
- Deep Repulsive Prototypes for Adversarial Robustness
- Model-Agnostic Meta-Attack: Towards Reliable Evaluation of Adversarial Robustness
- Adversarial Semantic Collisions
- Delving into Deep Image Prior for Adversarial Defense: A Novel Reconstruction-based Defense Framework
- A Neuro-Inspired Autoencoding Defense Against Adversarial Perturbations
- An Empirical Review of Adversarial Defenses
- Ensemble of Models Trained by Key-based Transformed Images for Adversarially Robust Defense Against Black-box Attacks
- Robust Vision-Based Cheat Detection in Competitive Gaming
- Sparse Coding Frontend for Robust Neural Networks
- Compressive Sensing Based Adaptive Defence Against Adversarial Images
- Tricking Adversarial Attacks To Fail
- Target Training Does Adversarial Training Without Adversarial Samples
- Adversarial Attacks with Time-Scale Representations
- Towards adversarial robustness with 01 loss neural networks
- A Closer Look at the Adversarial Robustness of Information Bottleneck Models
- Delving into the pixels of adversarial samples
- Attack to Fool and Explain Deep Networks
- Thwarting finite difference adversarial attacks with output randomization
- Constraining Logits by Bounded Function for Adversarial Robustness
- PDPGD: Primal-Dual Proximal Gradient Descent Adversarial Attack
- Hard-label Manifolds: Unexpected Advantages of Query Efficiency for Finding On-manifold Adversarial Examples