Fast is better than free: Revisiting adversarial training
arXiv:2001.03994
Abstract
Adversarial training, a method for learning robust deep networks, is typically assumed to be more expensive than traditional training due to the necessity of constructing adversarial examples via a first-order method like projected gradient decent (PGD). In this paper, we make the surprising discovery that it is possible to train empirically robust models using a much weaker and cheaper adversary, an approach that was previously believed to be ineffective, rendering the method no more costly than standard training in practice. Specifically, we show that adversarial training with the fast gradient sign method (FGSM), when combined with random initialization, is as effective as PGD-based training but has significantly lower cost. Furthermore we show that FGSM adversarial training can be further accelerated by using standard techniques for efficient training of deep networks, allowing us to learn a robust CIFAR10 classifier with 45% robust accuracy to PGD attacks with in 6 minutes, and a robust ImageNet classifier with 43% robust accuracy at in 12 hours, in comparison to past work based on "free" adversarial training which took 10 and 50 hours to reach the same respective thresholds. Finally, we identify a failure mode referred to as "catastrophic overfitting" which may have caused previous attempts to use FGSM adversarial training to fail. All code for reproducing the experiments in this paper as well as pretrained model weights are at https://github.com/locuslab/fast_adversarial.
References in corpus (7)
- Certified Adversarial Robustness via Randomized Smoothing
- On Evaluating Adversarial Robustness
- Countering Adversarial Images using Input Transformations
- Defensive Distillation is Not Robust to Adversarial Examples
- NO Need to Worry about Adversarial Examples in Object Detection in Autonomous Vehicles
- ME-Net: Towards Effective Adversarial Robustness with Matrix Estimation
- Is AmI (Attacks Meet Interpretability) Robust to Adversarial Examples?
Cited by in corpus (52)
- Invisible for both Camera and LiDAR: Security of Multi-Sensor Fusion based Perception in Autonomous Driving Under Physical-World Attacks
- Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead
- Overfitting in adversarially robust deep learning
- RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style Transformation
- Opportunities and Challenges in Deep Learning Adversarial Robustness: A Survey
- Exposing Previously Undetectable Faults in Deep Neural Networks
- Towards Understanding Fast Adversarial Training
- Dual Manifold Adversarial Robustness: Defense against Lp and non-Lp Adversarial Attacks
- This Looks Like That... Does it? Shortcomings of Latent Space Prototype Interpretability in Deep Networks
- On Fast Adversarial Robustness Adaptation in Model-Agnostic Meta-Learning
- A Real-time Defense against Website Fingerprinting Attacks
- Understanding the Interaction of Adversarial Training with Noisy Labels
- Attacking Adversarial Attacks as A Defense
- How benign is benign overfitting?
- DSRNA: Differentiable Search of Robust Neural Architectures
- Fighting Gradients with Gradients: Dynamic Defenses against Adversarial Attacks
- Adversarial Visual Robustness by Causal Intervention
- Guided Interpolation for Adversarial Training
- Adversarial Transfer Attacks With Unknown Data and Class Overlap
- AccelAT: A Framework for Accelerating the Adversarial Training of Deep Neural Networks through Accuracy Gradient
- Sequential Randomized Smoothing for Adversarially Robust Speech Recognition
- Amplification trojan network: Attack deep neural networks by amplifying their inherent weakness
- LAFEAT: Piercing Through Adversarial Defenses with Latent Features
- FAT: Federated Adversarial Training
- Effective, Efficient and Robust Neural Architecture Search
- Analysis and Applications of Class-wise Robustness in Adversarial Training
- Meta Adversarial Perturbations
- Anti-Bandit Neural Architecture Search for Model Defense
- Exploring Adversarial Robustness of Deep Metric Learning
- Multi-objective Search of Robust Neural Architectures against Multiple Types of Adversarial Attacks
- ZeroGrad : Mitigating and Explaining Catastrophic Overfitting in FGSM Adversarial Training
- Multi-stage Optimization based Adversarial Training
- Towards Speeding up Adversarial Training in Latent Spaces
- Adversarial Attacks on ML Defense Models Competition
- Model-Agnostic Meta-Attack: Towards Reliable Evaluation of Adversarial Robustness
- Improving Hierarchical Adversarial Robustness of Deep Neural Networks
- Utilizing Adversarial Targeted Attacks to Boost Adversarial Robustness
- "What's in the box?!": Deflecting Adversarial Attacks by Randomly Deploying Adversarially-Disjoint Models
- THAT: Two Head Adversarial Training for Improving Robustness at Scale
- Adaptive Clustering of Robust Semantic Representations for Adversarial Image Purification
- Deep Repulsive Prototypes for Adversarial Robustness
- Real-time Detection of Practical Universal Adversarial Perturbations
- Heating up decision boundaries: isocapacitory saturation, adversarial scenarios and generalization bounds
- Ensemble of Models Trained by Key-based Transformed Images for Adversarially Robust Defense Against Black-box Attacks
- Identifying and Exploiting Structures for Reliable Deep Learning
- Get Fooled for the Right Reason: Improving Adversarial Robustness through a Teacher-guided Curriculum Learning Approach
- Countering Adversarial Examples: Combining Input Transformation and Noisy Training
- Robust Face Verification via Disentangled Representations
- Reject Illegal Inputs with Generative Classifier Derived from Any Discriminative Classifier
- Skew Orthogonal Convolutions
- Black-box Adversarial Example Generation with Normalizing Flows
- Adversarial Training with Stochastic Weight Average