On the Convergence and Robustness of Adversarial Training
arXiv:2112.08304
Abstract
Improving the robustness of deep neural networks (DNNs) to adversarial examples is an important yet challenging problem for secure deep learning. Across existing defense techniques, adversarial training with Projected Gradient Decent (PGD) is amongst the most effective. Adversarial training solves a min-max optimization problem, with the \textit{inner maximization} generating adversarial examples by maximizing the classification loss, and the \textit{outer minimization} finding model parameters by minimizing the loss on adversarial examples generated from the inner maximization. A criterion that measures how well the inner maximization is solved is therefore crucial for adversarial training. In this paper, we propose such a criterion, namely First-Order Stationary Condition for constrained optimization (FOSC), to quantitatively evaluate the convergence quality of adversarial examples found in the inner maximization. With FOSC, we find that to ensure better robustness, it is essential to use adversarial examples with better convergence quality at the \textit{later stages} of training. Yet at the early stages, high convergence quality adversarial examples are not necessary and may even lead to poor robustness. Based on these observations, we propose a \textit{dynamic} training strategy to gradually increase the convergence quality of the generated adversarial examples, which significantly improves the robustness of adversarial training. Our theoretical and empirical results show the effectiveness of the proposed method.
ICML 2019 Long Talk. Fixing bugs in the proof of Theorem 1
References in corpus (4)
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Countering Adversarial Images using Input Transformations
- Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
Cited by in corpus (39)
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks
- Unlearnable Examples: Making Personal Data Unexploitable
- Recent Advances in Adversarial Training for Adversarial Robustness
- Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks
- PAD: Towards Principled Adversarial Malware Detection Against Evasion Attacks
- GANs May Have No Nash Equilibria
- Towards Understanding Fast Adversarial Training
- Adversarial training in communication constrained federated learning
- Understanding the Interaction of Adversarial Training with Noisy Labels
- Provably Efficient Black-Box Action Poisoning Attacks Against Reinforcement Learning
- Certified Robustness to Word Substitution Ranking Attack for Neural Ranking Models
- SoK: Machine Learning Governance
- Backdoor Attacks on Crowd Counting
- Black-box Adversarial Attacks on Video Recognition Models
- Adversarial Attacks on Black Box Video Classifiers: Leveraging the Power of Geometric Transformations
- Improving Query Efficiency of Black-box Adversarial Attack
- RayS: A Ray Searching Method for Hard-label Adversarial Attack
- Instance Correction for Learning with Open-set Noisy Labels
- Evaluating the Robustness of Geometry-Aware Instance-Reweighted Adversarial Training
- Guided Interpolation for Adversarial Training
- Revisiting Adversarial Robustness Distillation: Robust Soft Labels Make Student Better
- Local Reweighting for Adversarial Training
- Adversarially Robust Estimate and Risk Analysis in Linear Regression
- RobOT: Robustness-Oriented Testing for Deep Learning Systems
- Understanding the Intrinsic Robustness of Image Distributions using Conditional Generative Models
- Where is the Bottleneck of Adversarial Learning with Unlabeled Data?
- Analysis and Applications of Class-wise Robustness in Adversarial Training
- Adaptive Feature Alignment for Adversarial Training
- Multi-objective Search of Robust Neural Architectures against Multiple Types of Adversarial Attacks
- Enhance Diffusion to Improve Robust Generalization
- Non-Singular Adversarial Robustness of Neural Networks
- Provable Robustness of Adversarial Training for Learning Halfspaces with Noise
- Robust Deep Learning as Optimal Control: Insights and Convergence Guarantees
- ATRAS: Adversarially Trained Robust Architecture Search
- Model-Agnostic Defense for Lane Detection against Adversarial Attack
- Get Fooled for the Right Reason: Improving Adversarial Robustness through a Teacher-guided Curriculum Learning Approach
- Adversarial Interaction Attack: Fooling AI to Misinterpret Human Intentions
- Dual Head Adversarial Training
- Calibrated Adversarial Training