SafetyNet: Detecting and Rejecting Adversarial Examples Robustly
arXiv:1704.00103
Abstract
We describe a method to produce a network where current methods such as DeepFool have great difficulty producing adversarial samples. Our construction suggests some insights into how deep networks work. We provide a reasonable analyses that our construction is difficult to defeat, and show experimentally that our method is hard to defeat with both Type I and Type II attacks using several standard networks and datasets. This SafetyNet architecture is used to an important and novel application SceneProof, which can reliably detect whether an image is a picture of a real scene or not. SceneProof applies to images captured with depth maps (RGBD images) and checks if a pair of image and depth map is consistent. It relies on the relative difficulty of producing naturalistic depth maps for images in post processing. We demonstrate that our SafetyNet is robust to adversarial examples built from currently known attacking approaches.
Accepted to ICCV 2017
References in corpus (4)
Cited by in corpus (21)
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models
- Adversarial Examples that Fool Detectors
- EAD: Elastic-Net Attacks to Deep Neural Networks via Adversarial Examples
- Standard detectors aren't (currently) fooled by physical adversarial stop signs
- Daedalus: Breaking Non-Maximum Suppression in Object Detection via Adversarial Examples
- Classification regions of deep neural networks
- Defense against Universal Adversarial Perturbations
- Improved robustness to adversarial examples using Lipschitz regularization of the loss
- Improving Network Robustness against Adversarial Attacks with Compact Convolution
- The Taboo Trap: Behavioural Detection of Adversarial Samples
- Detecting Adversarial Examples in Convolutional Neural Networks
- Learning to Forecast and Refine Residual Motion for Image-to-Video Generation
- ReabsNet: Detecting and Revising Adversarial Examples
- Defending against Adversarial Attack towards Deep Neural Networks via Collaborative Multi-task Training
- Machine vs Machine: Minimax-Optimal Defense Against Adversarial Examples
- Butterfly Effect: Bidirectional Control of Classification Performance by Small Additive Perturbation
- Adversarial Open-World Person Re-Identification
- Efficient Diverse Ensemble for Discriminative Co-Tracking
- Utilizing Network Properties to Detect Erroneous Inputs
- Towards Robust Classification with Image Quality Assessment