Adversarial Machine Learning at Scale
arXiv:1611.01236
Abstract
Adversarial examples are malicious inputs designed to fool machine learning models. They often transfer from one model to another, allowing attackers to mount black box attacks without knowledge of the target model's parameters. Adversarial training is the process of explicitly training a model on adversarial examples, in order to make it more robust to attack or to reduce its test error on clean inputs. So far, adversarial training has primarily been applied to small problems. In this research, we apply adversarial training to ImageNet. Our contributions include: (1) recommendations for how to succesfully scale adversarial training to large models and datasets, (2) the observation that adversarial training confers robustness to single-step attack methods, (3) the finding that multi-step attack methods are somewhat less transferable than single-step attack methods, so single-step attacks are the best for mounting black-box attacks, and (4) resolution of a "label leaking" effect that causes adversarially trained models to perform better on adversarial examples than on clean examples, because the adversarial example construction process uses the true label and the model can learn to exploit regularities in the construction process.
17 pages, 5 figures
Cited by in corpus (50)
- Countering Adversarial Images using Input Transformations
- Measuring the tendency of CNNs to Learn Surface Statistical Regularities
- Fast Feature Fool: A data independent approach to universal adversarial perturbations
- Towards Interpretable Deep Neural Networks by Leveraging Adversarial Examples
- Optimization and Abstraction: A Synergistic Approach for Analyzing Neural Network Robustness
- Robust Android Malware Detection System against Adversarial Attacks using Q-Learning
- Improving the Transferability of Adversarial Examples with Resized-Diverse-Inputs, Diversity-Ensemble and Region Fitting
- Recent Advances in Adversarial Training for Adversarial Robustness
- Evading Defenses to Transferable Adversarial Examples by Translation-Invariant Attacks
- Detecting Adversarial Attacks on Neural Network Policies with Visual Foresight
- MaskDGA: A Black-box Evasion Technique Against DGA Classifiers and Adversarial Defenses
- A Survey of Black-Box Adversarial Attacks on Computer Vision Models
- GraphDefense: Towards Robust Graph Convolutional Networks
- Attacking Binarized Neural Networks
- There is Limited Correlation between Coverage and Robustness for Deep Neural Networks
- Risk Bounds for Robust Deep Learning
- AdvSPADE: Realistic Unrestricted Attacks for Semantic Segmentation
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- Better the Devil you Know: An Analysis of Evasion Attacks using Out-of-Distribution Adversarial Examples
- Second-Order Provable Defenses against Adversarial Attacks
- Adversarial Attack and Defense in Deep Ranking
- Dealing with Adversarial Player Strategies in the Neural Network Game iNNk through Ensemble Learning
- Federated AI lets a team imagine together: Federated Learning of GANs
- Robustness Certificates Against Adversarial Examples for ReLU Networks
- Differentiable Language Model Adversarial Attacks on Categorical Sequence Classifiers
- Input Hessian Regularization of Neural Networks
- Adversarial Defense Through Network Profiling Based Path Extraction
- Adaptive Feature Alignment for Adversarial Training
- Perturbations are not Enough: Generating Adversarial Examples with Spatial Distortions
- Assessing the Adversarial Robustness of Monte Carlo and Distillation Methods for Deep Bayesian Neural Network Classification
- Adversarial robustness via stochastic regularization of neural activation sensitivity
- Reporting on Decision-Making Algorithms and some Related Ethical Questions
- Deterministic Gaussian Averaged Neural Networks
- Provable Robustness of Adversarial Training for Learning Halfspaces with Noise
- Generating Unrestricted Adversarial Examples via Three Parameters
- The Vulnerability of the Neural Networks Against Adversarial Examples in Deep Learning Algorithms
- Non-Determinism in Neural Networks for Adversarial Robustness
- Extremal learning: extremizing the output of a neural network in regression problems
- Feature Losses for Adversarial Robustness
- Landmark Breaker: Obstructing DeepFake By Disturbing Landmark Extraction
- Polarizing Front Ends for Robust CNNs
- Robust Regularization with Adversarial Labelling of Perturbed Samples
- Recent Advancements in Self-Supervised Paradigms for Visual Feature Representation
- Multi-concept adversarial attacks
- On Connections between Regularizations for Improving DNN Robustness
- Countering Adversarial Examples: Combining Input Transformation and Noisy Training
- Constructing a provably adversarially-robust classifier from a high accuracy one
- GraCIAS: Grassmannian of Corrupted Images for Adversarial Security
- Strategies to architect AI Safety: Defense to guard AI from Adversaries
- Adversarial Attacks Against Deep Learning Systems for ICD-9 Code Assignment