RAB: Provable Robustness Against Backdoor Attacks
arXiv:2003.08904
Abstract
Recent studies have shown that deep neural networks (DNNs) are vulnerable to adversarial attacks, including evasion and backdoor (poisoning) attacks. On the defense side, there have been intensive efforts on improving both empirical and provable robustness against evasion attacks; however, the provable robustness against backdoor attacks still remains largely unexplored. In this paper, we focus on certifying the machine learning model robustness against general threat models, especially backdoor attacks. We first provide a unified framework via randomized smoothing techniques and show how it can be instantiated to certify the robustness against both evasion and backdoor attacks. We then propose the first robust training process, RAB, to smooth the trained model and certify its robustness against backdoor attacks. We prove the robustness bound for machine learning models trained with RAB and prove that our robustness bound is tight. In addition, we theoretically show that it is possible to train the robust smoothed models efficiently for simple models such as K-nearest neighbor classifiers, and we propose an exact smooth-training algorithm that eliminates the need to sample from a noise distribution for such models. Empirically, we conduct comprehensive experiments for different machine learning (ML) models such as DNNs, support vector machines, and K-NN models on MNIST, CIFAR-10, and ImageNette datasets and provide the first benchmark for certified robustness against backdoor attacks. In addition, we evaluate K-NN models on a spambase tabular dataset to demonstrate the advantages of the proposed exact algorithm. Both the theoretic analysis and the comprehensive evaluation on diverse ML models and datasets shed light on further robust learning strategies against general training time attacks.
IEEE Symposium on Security and Privacy 2023
References in corpus (8)
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Poisoning Attacks against Support Vector Machines
- An approach to reachability analysis for feed-forward ReLU neural networks
- Generative Poisoning Attack Method Against Neural Networks
- On Certifying Robustness against Backdoor Attacks via Randomized Smoothing
- Randomized Smoothing of All Shapes and Sizes
- Certified Robustness to Label-Flipping Attacks via Randomized Smoothing
- Improving Certified Robustness via Statistical Learning with Logical Reasoning
Cited by in corpus (19)
- Februus: Input Purification Defense Against Trojan Attacks on Deep Neural Network Systems
- Wild Patterns Reloaded: A Survey of Machine Learning Security against Training Data Poisoning
- Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review
- Backdoor Pre-trained Models Can Transfer to All
- Adversarial Unlearning of Backdoors via Implicit Hypergradient
- Adversarial Neuron Pruning Purifies Backdoored Deep Models
- TrojanZoo: Towards Unified, Holistic, and Practical Evaluation of Neural Backdoors
- DP-InstaHide: Provably Defusing Poisoning and Backdoor Attacks with Differentially Private Data Augmentations
- Deep Partition Aggregation: Provable Defense against General Poisoning Attacks
- SPECTRE: Defending Against Backdoor Attacks Using Robust Statistics
- SoK: Machine Learning Governance
- Improved, Deterministic Smoothing for L_1 Certified Robustness
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial Training
- Gradient Shaping: Enhancing Backdoor Attack Against Reverse Engineering
- On Convergence of Nearest Neighbor Classifiers over Feature Transformations
- A General Framework for Defending Against Backdoor Attacks via Influence Graph
- Learning and Certification under Instance-targeted Poisoning
- On Provable Backdoor Defense in Collaborative Learning
- A Framework of Randomized Selection Based Certified Defenses Against Data Poisoning Attacks