Clustering Effect of (Linearized) Adversarial Robust Models
arXiv:2111.12922
Abstract
Adversarial robustness has received increasing attention along with the study of adversarial examples. So far, existing works show that robust models not only obtain robustness against various adversarial attacks but also boost the performance in some downstream tasks. However, the underlying mechanism of adversarial robustness is still not clear. In this paper, we interpret adversarial robustness from the perspective of linear components, and find that there exist some statistical properties for comprehensively robust models. Specifically, robust models show obvious hierarchical clustering effect on their linearized sub-networks, when removing or replacing all non-linear components (e.g., batch normalization, maximum pooling, or activation layers). Based on these observations, we propose a novel understanding of adversarial robustness and apply it on more tasks including domain adaption and robustness boosting. Experimental evaluations demonstrate the rationality and superiority of our proposed clustering strategy.
Accepted by NeurIPS 2021, spotlight
References in corpus (9)
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
- Defensive Distillation is Not Robust to Adversarial Examples
- On the Convergence and Robustness of Adversarial Training
- Skip Connections Matter: On the Transferability of Adversarial Examples Generated with ResNets
- Residual Convolutional CTC Networks for Automatic Speech Recognition
- Improving Adversarial Robustness via Channel-wise Activation Suppressing
- Batch Normalization is a Cause of Adversarial Vulnerability
- BREEDS: Benchmarks for Subpopulation Shift
- Improve Generalization and Robustness of Neural Networks via Weight Scale Shifting Invariant Regularizations