Adversarial Weight Perturbation Helps Robust Generalization
arXiv:2004.05884
Abstract
The study on improving the robustness of deep neural networks against adversarial examples grows rapidly in recent years. Among them, adversarial training is the most promising one, which flattens the input loss landscape (loss change with respect to input) via training on adversarially perturbed examples. However, how the widely used weight loss landscape (loss change with respect to weight) performs in adversarial training is rarely explored. In this paper, we investigate the weight loss landscape from a new perspective, and identify a clear correlation between the flatness of weight loss landscape and robust generalization gap. Several well-recognized adversarial training improvements, such as early stopping, designing new objective functions, or leveraging unlabeled data, all implicitly flatten the weight loss landscape. Based on these observations, we propose a simple yet effective Adversarial Weight Perturbation (AWP) to explicitly regularize the flatness of weight loss landscape, forming a double-perturbation mechanism in the adversarial training framework that adversarially perturbs both inputs and weights. Extensive experiments demonstrate that AWP indeed brings flatter weight loss landscape and can be easily incorporated into various existing adversarial training methods to further boost their adversarial robustness.
To appear in NeurIPS 2020
References in corpus (13)
- Improved Regularization of Convolutional Neural Networks with Cutout
- Theoretically Principled Trade-off between Robustness and Accuracy
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
- Fast is better than free: Revisiting adversarial training
- Countering Adversarial Images using Input Transformations
- Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients
- On the Convergence and Robustness of Adversarial Training
- Skip Connections Matter: On the Transferability of Adversarial Examples Generated with ResNets
- Adversarial Robustness May Be at Odds With Simplicity
- Residual Convolutional CTC Networks for Automatic Speech Recognition
- VC Classes are Adversarially Robustly Learnable, but Only Improperly
- Understanding Adversarial Robustness Through Loss Landscape Geometries
Cited by in corpus (18)
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- A Review of Adversarial Attack and Defense for Classification Methods
- Semantic-Aware Adversarial Training for Reliable Deep Hashing Retrieval
- Noisy Differentiable Architecture Search
- Is Aggregation the Only Choice? Federated Learning via Layer-wise Model Recombination
- Adversarial Robustness via Fisher-Rao Regularization
- Efficient Sharpness-aware Minimization for Improved Training of Neural Networks
- SAT: Improving Adversarial Training via Curriculum-Based Loss Smoothing
- Computing-In-Memory Neural Network Accelerators for Safety-Critical Systems: Can Small Device Variations Be Disastrous?
- InfoAT: Improving Adversarial Training Using the Information Bottleneck Principle
- AdvFAS: A robust face anti-spoofing framework against adversarial examples
- Visualizing high-dimensional loss landscapes with Hessian directions
- BayesFT: Bayesian Optimization for Fault Tolerant Neural Network Architecture
- Towards Robust Neural Networks via Orthogonal Diversity
- Imbalanced Gradients: A Subtle Cause of Overestimated Adversarial Robustness
- Adversarial Robustness through the Lens of Convolutional Filters
- Bridging the Gap Between Adversarial Robustness and Optimization Bias
- MEAT: Median-Ensemble Adversarial Training for Improving Robustness and Generalization