Adversarial Examples in Modern Machine Learning: A Review
arXiv:1911.05268
Abstract
Recent research has found that many families of machine learning models are vulnerable to adversarial examples: inputs that are specifically designed to cause the target model to produce erroneous outputs. In this survey, we focus on machine learning models in the visual domain, where methods for generating and detecting such examples have been most extensively studied. We explore a variety of adversarial attack methods that apply to image-space content, real world adversarial attacks, adversarial defenses, and the transferability property of adversarial examples. We also discuss strengths and weaknesses of various methods of adversarial attack and defense. Our aim is to provide an extensive coverage of the field, furnishing the reader with an intuitive understanding of the mechanics of adversarial attack and defense mechanisms and enlarging the community of researchers studying this fundamental set of problems.
Work in progress, 97 pages
References in corpus (24)
- Distilling the Knowledge in a Neural Network
- Sequence to Sequence Learning with Neural Networks
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- On the difficulty of training Recurrent Neural Networks
- ZOO: Zeroth Order Optimization based Black-box Attacks to Deep Neural Networks without Training Substitute Models
- CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning
- Delving into Transferable Adversarial Examples and Black-box Attacks
- The Loss Surfaces of Multilayer Networks
- On Evaluating Adversarial Robustness
- The Space of Transferable Adversarial Examples
- A study of the effect of JPG compression on adversarial images
- Defensive Distillation is Not Robust to Adversarial Examples
- On Detecting Adversarial Perturbations
- Better Mixing via Deep Representations
- NO Need to Worry about Adversarial Examples in Object Detection in Autonomous Vehicles
- Did you hear that? Adversarial Examples Against Automatic Speech Recognition
- Measuring the tendency of CNNs to Learn Surface Statistical Regularities
- MagNet and "Efficient Defenses Against Adversarial Attacks" are Not Robust to Adversarial Examples
- Fast Feature Fool: A data independent approach to universal adversarial perturbations
- Blocking Transferability of Adversarial Examples in Black-Box Learning Systems
- Interpreting Adversarially Trained Convolutional Neural Networks
- Adversarial Attacks on Neural Network Policies
- Standard detectors aren't (currently) fooled by physical adversarial stop signs
- Transfer of Adversarial Robustness Between Perturbation Types