Adversarial Learning in Statistical Classification: A Comprehensive Review of Defenses Against Attacks
arXiv:1904.06292
Abstract
There is great potential for damage from adversarial learning (AL) attacks on machine-learning based systems. In this paper, we provide a contemporary survey of AL, focused particularly on defenses against attacks on statistical classifiers. After introducing relevant terminology and the goals and range of possible knowledge of both attackers and defenders, we survey recent work on test-time evasion (TTE), data poisoning (DP), and reverse engineering (RE) attacks and particularly defenses against same. In so doing, we distinguish robust classification from anomaly detection (AD), unsupervised from supervised, and statistical hypothesis-based defenses from ones that do not have an explicit null (no attack) hypothesis; we identify the hyperparameters a particular method requires, its computational complexity, as well as the performance measures on which it was evaluated and the obtained quality. We then dig deeper, providing novel insights that challenge conventional AL wisdom and that target unresolved issues, including: 1) robust classification versus AD as a defense strategy; 2) the belief that attack success increases with attack strength, which ignores susceptibility to AD; 3) small perturbations for test-time evasion attacks: a fallacy or a requirement?; 4) validity of the universal assumption that a TTE attacker knows the ground-truth class for the example to be attacked; 5) black, grey, or white box attacks as the standard for defense evaluation; 6) susceptibility of query-based RE to an AD defense. We also discuss attacks on the privacy of training data. We then present benchmark comparisons of several defenses against TTE, RE, and backdoor DP attacks on images. The paper concludes with a discussion of future work.
References in corpus (12)
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- On Evaluating Adversarial Robustness
- Security Evaluation of Pattern Classifiers under Attack
- Spectral Signatures in Backdoor Attacks
- Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
- Support Vector Machines under Adversarial Label Contamination
- On Detecting Adversarial Perturbations
- Early Methods for Detecting Adversarial Images
- Backdoor Embedding in Convolutional Neural Network Models via Invisible Perturbation
- Efficient and Accurate Estimation of Lipschitz Constants for Deep Neural Networks
- Revealing Perceptible Backdoors, without the Training Set, via the Maximum Achievable Misclassification Fraction Statistic
- A Mixture Model Based Defense for Data Poisoning Attacks Against Naive Bayes Spam Filters
Cited by in corpus (7)
- Recent advances for quantum classifiers
- Universal Adversarial Examples and Perturbations for Quantum Classifiers
- A Backdoor Attack against 3D Point Cloud Classifiers
- Revealing Perceptible Backdoors, without the Training Set, via the Maximum Achievable Misclassification Fraction Statistic
- Adversarial Example Detection by Classification for Deep Speech Recognition
- L-RED: Efficient Post-Training Detection of Imperceptible Backdoor Attacks without Access to the Training Set
- Reverse Engineering Imperceptible Backdoor Attacks on Deep Neural Networks for Detection and Training Set Cleansing