Improving Adversarial Robustness via Channel-wise Activation Suppressing
arXiv:2103.08307
Abstract
The study of adversarial examples and their activation has attracted significant attention for secure and robust learning with deep neural networks (DNNs). Different from existing works, in this paper, we highlight two new characteristics of adversarial examples from the channel-wise activation perspective: 1) the activation magnitudes of adversarial examples are higher than that of natural examples; and 2) the channels are activated more uniformly by adversarial examples than natural examples. We find that the state-of-the-art defense adversarial training has addressed the first issue of high activation magnitudes via training on adversarial examples, while the second issue of uniform activation remains. This motivates us to suppress redundant activation from being activated by adversarial perturbations via a Channel-wise Activation Suppressing (CAS) strategy. We show that CAS can train a model that inherently suppresses adversarial activation, and can be easily applied to existing defense methods to further improve their robustness. Our work provides a simple but generic training strategy for robustifying the intermediate layer activation of DNNs.
ICLR2021 accepted paper
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Theoretically Principled Trade-off between Robustness and Accuracy
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Residual Convolutional CTC Networks for Automatic Speech Recognition
- Batch Normalization is a Cause of Adversarial Vulnerability
- Overfitting in adversarially robust deep learning
- Harnessing the Vulnerability of Latent Layers in Adversarially Trained Models
Cited by in corpus (17)
- Towards Robust Neural Image Compression: Adversarial Attack and Model Finetuning
- Exploring Architectural Ingredients of Adversarially Robust Deep Neural Networks
- Reliable Adversarial Distillation with Unreliable Teachers
- Understanding the Interaction of Adversarial Training with Noisy Labels
- CIFS: Improving Adversarial Robustness of CNNs via Channel-wise Importance-based Feature Selection
- Adversarial Robustness through the Lens of Convolutional Filters
- Instance Correction for Learning with Open-set Noisy Labels
- Guided Interpolation for Adversarial Training
- Revisiting Adversarial Robustness Distillation: Robust Soft Labels Make Student Better
- Towards an Awareness of Time Series Anomaly Detection Models' Adversarial Vulnerability
- Learning Diverse-Structured Networks for Adversarial Robustness
- Analysis and Applications of Class-wise Robustness in Adversarial Training
- Clustering Effect of (Linearized) Adversarial Robust Models
- A Unified Game-Theoretic Interpretation of Adversarial Robustness
- Less is More: Feature Selection for Adversarial Robustness with Compressive Counter-Adversarial Attacks
- Identifying Layers Susceptible to Adversarial Attacks
- Dual Head Adversarial Training