TABOR: A Highly Accurate Approach to Inspecting and Restoring Trojan Backdoors in AI Systems
arXiv:1908.01763
Abstract
A trojan backdoor is a hidden pattern typically implanted in a deep neural network. It could be activated and thus forces that infected model behaving abnormally only when an input data sample with a particular trigger present is fed to that model. As such, given a deep neural network model and clean input samples, it is very challenging to inspect and determine the existence of a trojan backdoor. Recently, researchers design and develop several pioneering solutions to address this acute problem. They demonstrate the proposed techniques have a great potential in trojan detection. However, we show that none of these existing techniques completely address the problem. On the one hand, they mostly work under an unrealistic assumption (e.g. assuming availability of the contaminated training database). On the other hand, the proposed techniques cannot accurately detect the existence of trojan backdoors, nor restore high-fidelity trojan backdoor images, especially when the triggers pertaining to the trojan vary in size, shape and position. In this work, we propose TABOR, a new trojan detection technique. Conceptually, it formalizes a trojan detection task as a non-convex optimization problem, and the detection of a trojan backdoor as the task of resolving the optimization through an objective function. Different from the existing technique also modeling trojan detection as an optimization problem, TABOR designs a new objective function--under the guidance of explainable AI techniques as well as heuristics--that could guide optimization to identify a trojan backdoor in a more effective fashion. In addition, TABOR defines a new metric to measure the quality of a trojan backdoor identified. Using an anomaly detection method, we show the new metric could better facilitate TABOR to identify intentionally injected triggers in an infected model and filter out false alarms......
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Deep Learning Face Representation by Joint Identification-Verification
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- DeepID3: Face Recognition with Very Deep Neural Networks
- Understanding Impacts of High-Order Loss Approximations and Features in Deep Learning Interpretation
Cited by in corpus (46)
- Februus: Input Purification Defense Against Trojan Attacks on Deep Neural Network Systems
- Wild Patterns Reloaded: A Survey of Machine Learning Security against Training Data Poisoning
- Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review
- Poison Ink: Robust and Invisible Backdoor Attack
- Blind Backdoors in Deep Learning Models
- Rethinking the Trigger of Backdoor Attack
- NeuronInspect: Detecting Backdoors in Neural Networks via Output Explanations
- You Autocomplete Me: Poisoning Vulnerabilities in Neural Code Completion
- On Certifying Robustness against Backdoor Attacks via Randomized Smoothing
- Backdoor Scanning for Deep Neural Networks through K-Arm Optimization
- Invisible Backdoor Attacks on Deep Neural Networks via Steganography and Regularization
- Adversarial Unlearning of Backdoors via Implicit Hypergradient
- DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
- Backdoor Attack through Frequency Domain
- TrojanZoo: Towards Unified, Holistic, and Practical Evaluation of Neural Backdoors
- Practical Detection of Trojan Neural Networks: Data-Limited and Data-Free Cases
- AEVA: Black-box Backdoor Detection Using Adversarial Extreme Value Analysis
- Selective Amnesia: On Efficient, High-Fidelity and Blind Suppression of Backdoor Effects in Trojaned Machine Learning Models
- Rethinking the Backdoor Attacks' Triggers: A Frequency Perspective
- Design and Evaluation of a Multi-Domain Trojan Detection Method on Deep Neural Networks
- Poisoned classifiers are not only backdoored, they are fundamentally broken
- Detecting Backdoors in Neural Networks Using Novel Feature-Based Anomaly Detection
- T-Miner: A Generative Approach to Defend Against Trojan Attacks on DNN-based Text Classification
- An Embarrassingly Simple Approach for Trojan Attack in Deep Neural Networks
- Backdoor Attacks to Graph Neural Networks
- TOP: Backdoor Detection in Neural Networks via Transferability of Perturbation
- Defending Against Backdoor Attacks in Natural Language Generation
- One-pixel Signature: Characterizing CNN Models for Backdoor Detection
- TAD: Trigger Approximation based Black-box Trojan Detection for AI
- New Security Challenges on Machine Learning Inference Engine: Chip Cloning and Model Reverse Engineering
- Trigger Hunting with a Topological Prior for Trojan Detection
- BACKDOORL: Backdoor Attack against Competitive Reinforcement Learning
- EX-RAY: Distinguishing Injected Backdoor from Natural Features in Neural Networks by Examining Differential Feature Symmetry
- A Backdoor Attack against 3D Point Cloud Classifiers
- PiDAn: A Coherence Optimization Approach for Backdoor Attack Detection and Mitigation in Deep Neural Networks
- Revealing Perceptible Backdoors, without the Training Set, via the Maximum Achievable Misclassification Fraction Statistic
- Cassandra: Detecting Trojaned Networks from Adversarial Perturbations
- Gradient Shaping: Enhancing Backdoor Attack Against Reverse Engineering
- Hidden Backdoors in Human-Centric Language Models
- NTD: Non-Transferability Enabled Backdoor Detection
- Topological Detection of Trojaned Neural Networks
- Towards Practical Deployment-Stage Backdoor Attack on Deep Neural Networks
- SGBA: A Stealthy Scapegoat Backdoor Attack against Deep Neural Networks
- PoisHygiene: Detecting and Mitigating Poisoning Attacks in Neural Networks
- Reverse Engineering Imperceptible Backdoor Attacks on Deep Neural Networks for Detection and Training Set Cleansing
- L-RED: Efficient Post-Training Detection of Imperceptible Backdoor Attacks without Access to the Training Set