On Evaluating Adversarial Robustness
arXiv:1902.06705
Abstract
Correctly evaluating defenses against adversarial examples has proven to be extremely difficult. Despite the significant amount of recent work attempting to design defenses that withstand adaptive attacks, few have succeeded; most papers that propose defenses are quickly shown to be incorrect. We believe a large contributing factor is the difficulty of performing security evaluations. In this paper, we discuss the methodological foundations, review commonly accepted best practices, and suggest new methods for evaluating defenses to adversarial examples. We hope that both researchers developing defenses as well as readers and reviewers who wish to understand the completeness of an evaluation consider our advice in order to avoid common pitfalls.
Living document; source available at https://github.com/evaluating-adversarial-robustness/adv-eval-paper/
References in corpus (8)
- Explaining and Harnessing Adversarial Examples
- ZOO: Zeroth Order Optimization based Black-box Attacks to Deep Neural Networks without Training Substitute Models
- Delving into Transferable Adversarial Examples and Black-box Attacks
- Defensive Distillation is Not Robust to Adversarial Examples
- MagNet and "Efficient Defenses Against Adversarial Attacks" are Not Robust to Adversarial Examples
- Adversarial Examples Are a Natural Consequence of Test Error in Noise
- Adversarial Example Defenses: Ensembles of Weak Defenses are not Strong
- Is AmI (Attacks Meet Interpretability) Robust to Adversarial Examples?
Cited by in corpus (11)
- Theoretical evidence for adversarial robustness through randomization
- Transfer of Adversarial Robustness Between Perturbation Types
- Exploiting Excessive Invariance caused by Norm-Bounded Adversarial Robustness
- A critique of the DeepSec Platform for Security Analysis of Deep Learning Models
- Adversarial Security Attacks and Perturbations on Machine Learning and Deep Learning Methods
- Interpreting Adversarial Examples with Attributes
- Stateful Detection of Black-Box Adversarial Attacks
- Bandlimiting Neural Networks Against Adversarial Attacks
- A unified view on differential privacy and robustness to adversarial examples
- Moving Target Defense for Deep Visual Sensing against Adversarial Examples
- Thwarting finite difference adversarial attacks with output randomization