Defense Against Adversarial Attacks Using Feature Scattering-based Adversarial Training
arXiv:1907.10764
Abstract
We introduce a feature scattering-based adversarial training approach for improving model robustness against adversarial attacks. Conventional adversarial training approaches leverage a supervised scheme (either targeted or non-targeted) in generating attacks for training, which typically suffer from issues such as label leaking as noted in recent works. Differently, the proposed approach generates adversarial images for training through feature scattering in the latent space, which is unsupervised in nature and avoids label leaking. More importantly, this new approach generates perturbed images in a collaborative fashion, taking the inter-sample relationships into consideration. We conduct analysis on model robustness and demonstrate the effectiveness of the proposed approach through extensively experiments on different datasets compared with state-of-the-art approaches.
Published at NeurIPS 2019
References in corpus (21)
- Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning
- Adversarially Robust Generalization Requires More Data
- The Space of Transferable Adversarial Examples
- Adversarial Examples Are Not Bugs, They Are Features
- PixelDefend: Leveraging Generative Models to Understand and Defend against Adversarial Examples
- Exploring the Landscape of Spatial Robustness
- Generating Adversarial Examples with Adversarial Networks
- On Detecting Adversarial Perturbations
- Essentially No Barriers in Neural Network Energy Landscape
- Houdini: Fooling Deep Structured Prediction Models
- A Boundary Tilting Persepective on the Phenomenon of Adversarial Examples
- The Robust Manifold Defense: Adversarial Training using Generative Models
- Physical Adversarial Examples for Object Detectors
- Deep Defense: Training DNNs with Improved Adversarial Robustness
- Feature Denoising for Improving Adversarial Robustness
- Adversarial Text Generation via Feature-Mover's Distance
- The Limitations of Adversarial Training and the Blind-Spot Attack
- GAN and VAE from an Optimal Transport Point of View
- On the Connection Between Adversarial Robustness and Saliency Map Interpretability
- Joint Adversarial Training: Incorporating both Spatial and Pixel Attacks
- AutoGAN: Robust Classifier Against Adversarial Attacks
Cited by in corpus (15)
- Adversarial Weight Perturbation Helps Robust Generalization
- Improving deep learning with prior knowledge and cognitive models: A survey on enhancing explainability, adversarial robustness and zero-shot learning
- Interpolated Adversarial Training: Achieving Robust Neural Networks without Sacrificing Too Much Accuracy
- Predify: Augmenting deep neural networks with brain-inspired predictive coding dynamics
- Black-Box Certification with Randomized Smoothing: A Functional Optimization Based Framework
- Adversarial Reinforced Instruction Attacker for Robust Vision-Language Navigation
- InfoAT: Improving Adversarial Training Using the Information Bottleneck Principle
- Jacobian Adversarially Regularized Networks for Robustness
- Imbalanced Gradients: A Subtle Cause of Overestimated Adversarial Robustness
- Defending against Adversarial Attack towards Deep Neural Networks via Collaborative Multi-task Training
- Towards an Awareness of Time Series Anomaly Detection Models' Adversarial Vulnerability
- A Self-supervised Approach for Adversarial Robustness
- What it Thinks is Important is Important: Robustness Transfers through Input Gradients
- Improving White-box Robustness of Pre-processing Defenses via Joint Adversarial Training
- Enhance Diffusion to Improve Robust Generalization