A Unified Approach to Interpreting and Boosting Adversarial Transferability
arXiv:2010.04055
Abstract
In this paper, we use the interaction inside adversarial perturbations to explain and boost the adversarial transferability. We discover and prove the negative correlation between the adversarial transferability and the interaction inside adversarial perturbations. The negative correlation is further verified through different DNNs with various inputs. Moreover, this negative correlation can be regarded as a unified perspective to understand current transferability-boosting methods. To this end, we prove that some classic methods of enhancing the transferability essentially decease interactions inside adversarial perturbations. Based on this, we propose to directly penalize interactions during the attacking process, which significantly improves the adversarial transferability.
References in corpus (6)
- ZOO: Zeroth Order Optimization based Black-box Attacks to Deep Neural Networks without Training Substitute Models
- Delving into Transferable Adversarial Examples and Black-box Attacks
- Skip Connections Matter: On the Transferability of Adversarial Examples Generated with ResNets
- Enhancing Adversarial Example Transferability with an Intermediate Level Attack
- Transferable Perturbations of Deep Feature Distributions
- Interpreting and Boosting Dropout from a Game-Theoretic View