Improving Adversarial Robustness of Ensembles with Diversity Training
arXiv:1901.09981
Abstract
Deep Neural Networks are vulnerable to adversarial attacks even in settings where the attacker has no direct access to the model being attacked. Such attacks usually rely on the principle of transferability, whereby an attack crafted on a surrogate model tends to transfer to the target model. We show that an ensemble of models with misaligned loss gradients can provide an effective defense against transfer-based attacks. Our key insight is that an adversarial example is less likely to fool multiple models in the ensemble if their loss functions do not increase in a correlated fashion. To this end, we propose Diversity Training, a novel method to train an ensemble of models with uncorrelated loss functions. We show that our method significantly improves the adversarial robustness of ensembles and can also be combined with existing methods to create a stronger defense.
References in corpus (2)
Cited by in corpus (12)
- Recent Advances in Adversarial Training for Adversarial Robustness
- Enhancing Certifiable Robustness via a Deep Model Ensemble
- DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial Estimation
- -ML: Mitigating Adversarial Examples via Ensembles of Topologically Manipulated Classifiers
- Ensemble Defense with Data Diversity: Weak Correlation Implies Strong Robustness
- Ensemble-in-One: Learning Ensemble within Random Gated Networks for Enhanced Adversarial Robustness
- Robustness from Simple Classifiers
- "What's in the box?!": Deflecting Adversarial Attacks by Randomly Deploying Adversarially-Disjoint Models
- Evaluating Ensemble Robustness Against Adversarial Attacks
- Orthogonal Deep Models As Defense Against Black-Box Attacks
- CoG: a Two-View Co-training Framework for Defending Adversarial Attacks on Graph
- Adversarial Training with Stochastic Weight Average