Fighting Gradients with Gradients: Dynamic Defenses against Adversarial Attacks
arXiv:2105.08714
Abstract
Adversarial attacks optimize against models to defeat defenses. Existing defenses are static, and stay the same once trained, even while attacks change. We argue that models should fight back, and optimize their defenses against attacks at test time. We propose dynamic defenses, to adapt the model and input during testing, by defensive entropy minimization (dent). Dent alters testing, but not training, for compatibility with existing models and train-time defenses. Dent improves the robustness of adversarially-trained defenses and nominally-trained models against white-box, black-box, and adaptive attacks on CIFAR-10/100 and ImageNet. In particular, dent boosts state-of-the-art defenses by 20+ points absolute against AutoAttack on CIFAR-10 at = 8/255.
References in corpus (11)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Theoretically Principled Trade-off between Robustness and Accuracy
- Certified Adversarial Robustness via Randomized Smoothing
- Fast is better than free: Revisiting adversarial training
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Deep Equilibrium Models
- Adversarial Examples Are a Natural Consequence of Test Error in Noise
- An Adaptive and Momental Bound Method for Stochastic Learning
- Blurring the Line Between Structure and Learning to Optimize and Adapt Receptive Fields
- A Research Agenda: Dynamic Models to Defend Against Correlated Attacks