Characterizing Adversarial Subspaces Using Local Intrinsic Dimensionality
arXiv:1801.02613
Abstract
Deep Neural Networks (DNNs) have recently been shown to be vulnerable against adversarial examples, which are carefully crafted instances that can mislead DNNs to make errors during prediction. To better understand such attacks, a characterization is needed of the properties of regions (the so-called 'adversarial subspaces') in which adversarial examples lie. We tackle this challenge by characterizing the dimensional properties of adversarial regions, via the use of Local Intrinsic Dimensionality (LID). LID assesses the space-filling capability of the region surrounding a reference example, based on the distance distribution of the example to its neighbors. We first provide explanations about how adversarial perturbation can affect the LID characteristic of adversarial regions, and then show empirically that LID characteristics can facilitate the distinction of adversarial examples generated using state-of-the-art attacks. As a proof-of-concept, we show that a potential application of LID is to distinguish adversarial examples, and the preliminary results show that it can outperform several state-of-the-art detection measures by large margins for five attack strategies considered in this paper across three benchmark datasets. Our analysis of the LID characteristic for adversarial regions not only motivates new directions of effective adversarial defense, but also opens up more challenges for developing new attacks to better understand the vulnerabilities of DNNs.
References in corpus (5)
- Delving into Transferable Adversarial Examples and Black-box Attacks
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- The Space of Transferable Adversarial Examples
- A Boundary Tilting Persepective on the Phenomenon of Adversarial Examples
- Adversarial Example Defenses: Ensembles of Weak Defenses are not Strong
Cited by in corpus (112)
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
- Theoretically Principled Trade-off between Robustness and Accuracy
- Understanding Adversarial Attacks on Deep Learning Based Medical Image Analysis Systems
- Guiding Deep Learning System Testing using Surprise Adequacy
- Adversarial Examples: Opportunities and Challenges
- Dimensionality-Driven Learning with Noisy Labels
- Adversarial Weight Perturbation Helps Robust Generalization
- Adversarial Example Detection for DNN Models: A Review and Experimental Comparison
- Adversarial Machine Learning in Image Classification: A Survey Towards the Defender's Perspective
- Motivating the Rules of the Game for Adversarial Example Research
- Intrinsic dimension of data representations in deep neural networks
- advertorch v0.1: An Adversarial Robustness Toolbox based on PyTorch
- Interpreting and Improving Adversarial Robustness of Deep Neural Networks with Neuron Sensitivity
- A Review of Adversarial Attack and Defense for Classification Methods
- AttriGuard: A Practical Defense Against Attribute Inference Attacks via Adversarial Machine Learning
- DeepRobust: A PyTorch Library for Adversarial Attacks and Defenses
- Gotta Catch 'Em All: Using Honeypots to Catch Adversarial Attacks on Neural Networks
- Adversarial Attack and Defense on Point Sets
- The Limitations of Adversarial Training and the Blind-Spot Attack
- Security and Privacy Issues in Deep Learning
- Detection of Face Recognition Adversarial Attacks
- Convergence of Adversarial Training in Overparametrized Neural Networks
- A Survey of Safety and Trustworthiness of Deep Neural Networks: Verification, Testing, Adversarial Attack and Defence, and Interpretability
- RAB: Provable Robustness Against Backdoor Attacks
- Towards Robust Neural Networks via Random Self-ensemble
- NATTACK: Learning the Distributions of Adversarial Examples for an Improved Black-Box Attack on Deep Neural Networks
- Towards Stable and Efficient Training of Verifiably Robust Neural Networks
- Confidence-Calibrated Adversarial Training: Generalizing to Unseen Attacks
- Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models
- Daedalus: Breaking Non-Maximum Suppression in Object Detection via Adversarial Examples
- Toward Robust Image Classification
- Exploring Architectural Ingredients of Adversarially Robust Deep Neural Networks
- On the Limitation of Local Intrinsic Dimensionality for Characterizing the Subspaces of Adversarial Examples
- Adversarial Examples in Deep Learning: Characterization and Divergence
- Understanding the One-Pixel Attack: Propagation Maps and Locality Analysis
- High Dimensional Spaces, Deep Learning and Adversarial Examples
- AdvFlow: Inconspicuous Black-box Adversarial Attacks using Normalizing Flows
- A Simple Fine-tuning Is All You Need: Towards Robust Deep Learning Via Adversarial Fine-tuning
- Evolving Robust Neural Architectures to Defend from Adversarial Attacks
- Adversarial Camouflage: Hiding Physical-World Attacks with Natural Styles
- Regional Homogeneity: Towards Learning Transferable Universal Adversarial Perturbations Against Defenses
- Defending Against Adversarial Examples with K-Nearest Neighbor
- Clean-Label Backdoor Attacks on Video Recognition Models
- Interpretable Deep Learning under Fire
- Enhancing the Robustness of Deep Neural Networks by Boundary Conditional GAN
- Breaking Transferability of Adversarial Samples with Randomness
- DaST: Data-free Substitute Training for Adversarial Attacks
- Adversarial Image Color Transformations in Explicit Color Filter Space
- Bag of Tricks for Adversarial Training
- Towards Security Threats of Deep Learning Systems: A Survey
- Enhancing Gradient-based Attacks with Symbolic Intervals
- Detecting Adversarial Examples Is (Nearly) As Hard As Classifying Them
- Distributionally Adversarial Attack
- On Certifying Non-uniform Bound against Adversarial Attacks
- Adversarial Robustness Assessment: Why both and Attacks Are Necessary
- Anomalous Example Detection in Deep Learning: A Survey
- Denoised Internal Models: a Brain-Inspired Autoencoder against Adversarial Attacks
- Using Undervolting as an On-Device Defense Against Adversarial Machine Learning Attacks
- Effective and Robust Detection of Adversarial Examples via Benford-Fourier Coefficients
- On the Certified Robustness for Ensemble Models and Beyond
- Detecting Adversarial Examples via Neural Fingerprinting
- LiBRe: A Practical Bayesian Approach to Adversarial Detection
- Instance Correction for Learning with Open-set Noisy Labels
- Imbalanced Gradients: A Subtle Cause of Overestimated Adversarial Robustness
- Defending against Adversarial Attack towards Deep Neural Networks via Collaborative Multi-task Training
- TSS: Transformation-Specific Smoothing for Robustness Certification
- Attacking and Defending Machine Learning Applications of Public Cloud
- Enhancing Adversarial Defense by k-Winners-Take-All
- Adversarial Defense by Stratified Convolutional Sparse Coding
- Almost Tight L0-norm Certified Robustness of Top-k Predictions against Adversarial Perturbations
- Evading Adversarial Example Detection Defenses with Orthogonal Projected Gradient Descent
- Random Directional Attack for Fooling Deep Neural Networks
- Defending against adversarial attacks on medical imaging AI system, classification or detection?
- The Adversarial Attack and Detection under the Fisher Information Metric
- RobOT: Robustness-Oriented Testing for Deep Learning Systems
- One Bit Matters: Understanding Adversarial Examples as the Abuse of Redundancy
- Customizing an Adversarial Example Generator with Class-Conditional GANs
- Trust but Verify: An Information-Theoretic Explanation for the Adversarial Fragility of Machine Learning Systems, and a General Defense against Adversarial Attacks
- Are Adversarial Examples Created Equal? A Learnable Weighted Minimax Risk for Robustness under Non-uniform Attacks
- Toward Metrics for Differentiating Out-of-Distribution Sets
- Intrinsic Geometric Vulnerability of High-Dimensional Artificial Intelligence
- Purifying Adversarial Perturbation with Adversarially Trained Auto-encoders
- Medical Aegis: Robust adversarial protectors for medical images
- Uncovering the Connections Between Adversarial Transferability and Knowledge Transferability
- A Unified Game-Theoretic Interpretation of Adversarial Robustness
- Detecting Patch Adversarial Attacks with Image Residuals
- Interpretable Disentanglement of Neural Networks by Extracting Class-Specific Subnetwork
- Improving White-box Robustness of Pre-processing Defenses via Joint Adversarial Training
- On The Utility of Conditional Generation Based Mutual Information for Characterizing Adversarial Subspaces
- Representation Quality Of Neural Networks Links To Adversarial Attacks and Defences
- Perceptual Deep Neural Networks: Adversarial Robustness through Input Recreation
- Adv-4-Adv: Thwarting Changing Adversarial Perturbations via Adversarial Domain Adaptation
- Exploiting the Sensitivity of Adversarial Examples to Erase-and-Restore
- Multi-Expert Adversarial Attack Detection in Person Re-identification Using Context Inconsistency
- A Hierarchical Feature Constraint to Camouflage Medical Adversarial Attacks
- MixDefense: A Defense-in-Depth Framework for Adversarial Example Detection Based on Statistical and Semantic Analysis
- A Survey on Assessing the Generalization Envelope of Deep Neural Networks: Predictive Uncertainty, Out-of-distribution and Adversarial Samples
- Adversarial Examples Detection beyond Image Space
- The art of defense: letting networks fool the attacker
- Training Provably Robust Models by Polyhedral Envelope Regularization
- Weighted Average Precision: Adversarial Example Detection in the Visual Perception of Autonomous Vehicles
- Adversarial Imitation Attack
- Defending SVMs against Poisoning Attacks: the Hardness and DBSCAN Approach
- Likelihood Landscapes: A Unifying Principle Behind Many Adversarial Defenses
- On Configurable Defense against Adversarial Example Attacks
- Learning To Characterize Adversarial Subspaces
- Linking average- and worst-case perturbation robustness via class selectivity and dimensionality
- A Sweet Rabbit Hole by DARCY: Using Honeypots to Detect Universal Trigger's Adversarial Attacks
- Generative Image Inpainting with Submanifold Alignment
- Denoising and Verification Cross-Layer Ensemble Against Black-box Adversarial Attacks
- BAARD: Blocking Adversarial Examples by Testing for Applicability, Reliability and Decidability
- Improving Adversarial Robustness for Free with Snapshot Ensemble