Gradient Masking Causes CLEVER to Overestimate Adversarial Perturbation Size
arXiv:1804.07870
Abstract
A key problem in research on adversarial examples is that vulnerability to adversarial examples is usually measured by running attack algorithms. Because the attack algorithms are not optimal, the attack algorithms are prone to overestimating the size of perturbation needed to fool the target model. In other words, the attack-based methodology provides an upper-bound on the size of a perturbation that will fool the model, but security guarantees require a lower bound. CLEVER is a proposed scoring method to estimate a lower bound. Unfortunately, an estimate of a bound is not a bound. In this report, we show that gradient masking, a common problem that causes attack methodologies to provide only a very loose upper bound, causes CLEVER to overestimate the size of perturbation needed to fool the model. In other words, CLEVER does not resolve the key problem with the attack-based methodology, because it fails to provide a lower bound.
References in corpus (2)
Cited by in corpus (9)
- Testing Robustness Against Unforeseen Adversaries
- A Survey of Safety and Trustworthiness of Deep Neural Networks: Verification, Testing, Adversarial Attack and Defence, and Interpretability
- A Statistical Approach to Assessing Neural Network Robustness
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- Robustness Certificates Against Adversarial Examples for ReLU Networks
- Global Adversarial Attacks for Assessing Deep Learning Robustness
- On Extensions of CLEVER: A Neural Network Robustness Evaluation Algorithm
- Toward Few-step Adversarial Training from a Frequency Perspective
- An Efficient and Margin-Approaching Zero-Confidence Adversarial Attack