Certified Defenses against Adversarial Examples
arXiv:1801.09344
Abstract
While neural networks have achieved high accuracy on standard image classification benchmarks, their accuracy drops to nearly zero in the presence of small adversarial perturbations to test inputs. Defenses based on regularization and adversarial training have been proposed, but often followed by new, stronger attacks that defeat these defenses. Can we somehow end this arms race? In this work, we study this problem for neural networks with one hidden layer. We first propose a method based on a semidefinite relaxation that outputs a certificate that for a given network and test input, no attack can force the error to exceed a certain value. Second, as this certificate is differentiable, we jointly optimize it with the network parameters, providing an adaptive regularizer that encourages robustness against all attacks. On MNIST, our approach produces a network and a certificate that no attack that perturbs each pixel by at most ε= 0.1 can cause more than 35% test error.
Published at the International Conference on Learning Representations (ICLR) 2018
References in corpus (4)
Cited by in corpus (199)
- On Evaluating Adversarial Robustness
- Adversarially Robust Generalization Requires More Data
- Adversarial Examples Are Not Bugs, They Are Features
- Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning
- Robustness May Be at Odds with Accuracy
- Efficient Neural Network Robustness Certification with General Activation Functions
- Certifying Some Distributional Robustness with Principled Adversarial Training
- On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models
- Provably Robust Deep Learning via Adversarially Trained Smoothed Classifiers
- Security and Privacy Approaches in Mixed Reality: A Literature Survey
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- Motivating the Rules of the Game for Adversarial Example Research
- Algorithmic decision-making in AVs: Understanding ethical and technical concerns for smart cities
- Adversarial Training and Robustness for Multiple Perturbations
- Constructing Unrestricted Adversarial Examples with Generative Models
- Rademacher Complexity for Adversarially Robust Generalization
- Do Adversarially Robust ImageNet Models Transfer Better?
- Adversarial Attacks Against Medical Deep Learning Systems
- Blind Backdoors in Deep Learning Models
- Training for Faster Adversarial Robustness Verification via Inducing ReLU Stability
- Unsolved Problems in ML Safety
- Towards quantum enhanced adversarial robustness in machine learning
- On the Robustness of Interpretability Methods
- Test-Time Training with Self-Supervision for Generalization under Distribution Shifts
- Unlabeled Data Improves Adversarial Robustness
- On the (In)fidelity and Sensitivity for Explanations
- Improving the Transferability of Adversarial Examples with Resized-Diverse-Inputs, Diversity-Ensemble and Region Fitting
- The Limitations of Adversarial Training and the Blind-Spot Attack
- Security and Privacy Issues in Deep Learning
- Excessive Invariance Causes Adversarial Vulnerability
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial Perturbations
- A Survey of Safety and Trustworthiness of Deep Neural Networks: Verification, Testing, Adversarial Attack and Defence, and Interpretability
- Adversarial Robustness of Deep Neural Networks: A Survey from a Formal Verification Perspective
- Tight Certificates of Adversarial Robustness for Randomly Smoothed Classifiers
- RAB: Provable Robustness Against Backdoor Attacks
- Certified Adversarial Robustness for Deep Reinforcement Learning
- Efficient and Accurate Estimation of Lipschitz Constants for Deep Neural Networks
- Diversity can be Transferred: Output Diversification for White- and Black-box Attacks
- Overfitting in adversarially robust deep learning
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- Understanding and Improving Fast Adversarial Training
- Logit Pairing Methods Can Fool Gradient-Based Attacks
- VC Classes are Adversarially Robustly Learnable, but Only Improperly
- Learning perturbation sets for robust machine learning
- Towards Stable and Efficient Training of Verifiably Robust Neural Networks
- Denoised Smoothing: A Provable Defense for Pretrained Classifiers
- Provably Robust Boosted Decision Stumps and Trees against Adversarial Attacks
- PatchGuard: A Provably Robust Defense against Adversarial Patches via Small Receptive Fields and Masking
- On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome Them
- DNNV: A Framework for Deep Neural Network Verification
- Algorithms for Verifying Deep Neural Networks
- Adversarial Robustness for Code
- Defend Deep Neural Networks Against Adversarial Examples via Fixed and Dynamic Quantized Activation Functions
- Towards a Robust and Trustworthy Machine Learning System Development: An Engineering Perspective
- L2-Nonexpansive Neural Networks
- Randomized Smoothing of All Shapes and Sizes
- Graph Adversarial Training: Dynamically Regularizing Based on Graph Structure
- Transferable Adversarial Attacks for Image and Video Object Detection
- Globally-Robust Neural Networks
- Adversarial Robustness via Fisher-Rao Regularization
- Data Poisoning against Differentially-Private Learners: Attacks and Defenses
- Learning Security Classifiers with Verified Global Robustness Properties
- Dirty Road Can Attack: Security of Deep Learning based Automated Lane Centering under Physical-World Attack
- Towards Understanding Fast Adversarial Training
- ADef: an Iterative Algorithm to Construct Adversarial Deformations
- Adversarial Risk Bounds via Function Transformation
- Adversarial Learning in Statistical Classification: A Comprehensive Review of Defenses Against Attacks
- Viewmaker Networks: Learning Views for Unsupervised Representation Learning
- Fast Neural Network Verification via Shadow Prices
- Morphence: Moving Target Defense Against Adversarial Examples
- Provable Certificates for Adversarial Examples: Fitting a Ball in the Union of Polytopes
- PatchCleanser: Certifiably Robust Defense against Adversarial Patches for Any Image Classifier
- More Data Can Expand the Generalization Gap Between Adversarially Robust and Standard Models
- Model Assertions for Monitoring and Improving ML Models
- Achieving Verified Robustness to Symbol Substitutions via Interval Bound Propagation
- Improving adversarial robustness of deep neural networks by using semantic information
- Adversarial Machine Learning in Wireless Communications using RF Data: A Review
- Semialgebraic Optimization for Lipschitz Constants of ReLU Networks
- Certified Robustness to Adversarial Word Substitutions
- Physically Realizable Adversarial Examples for LiDAR Object Detection
- A Provable Defense for Deep Residual Networks
- Robust Deep Reinforcement Learning through Adversarial Loss
- ResNets Ensemble via the Feynman-Kac Formalism to Improve Natural and Robust Accuracies
- Adversarial Framework with Certified Robustness for Time-Series Domain via Statistical Features
- A Spectral View of Adversarially Robust Features
- Towards Understanding the Adversarial Vulnerability of Skeleton-based Action Recognition
- Utility is in the Eye of the User: A Critique of NLP Leaderboards
- Robustifying Models Against Adversarial Attacks by Langevin Dynamics
- Defending Against Physically Realizable Attacks on Image Classification
- Detecting Adversarial Examples Is (Nearly) As Hard As Classifying Them
- DeepCover: Advancing RNN Test Coverage and Online Error Prediction using State Machine Extraction
- SoK: Machine Learning Governance
- Counterexample-Guided Learning of Monotonic Neural Networks
- Calibrated Surrogate Losses for Adversarially Robust Classification
- Enhancing Certifiable Robustness via a Deep Model Ensemble
- Understanding Generalization in Adversarial Training via the Bias-Variance Decomposition
- Distributionally Robust Local Non-parametric Conditional Estimation
- Efficient Adversarial Training with Transferable Adversarial Examples
- Universal Lipschitz Approximation in Bounded Depth Neural Networks
- Correctness Verification of Neural Networks
- Towards neural networks that provably know when they don't know
- Improved Image Wasserstein Attacks and Defenses
- GhostImage: Remote Perception Attacks against Camera-based Image Classification Systems
- On Certifying Non-uniform Bound against Adversarial Attacks
- Understanding Adversarial Robustness: The Trade-off between Minimum and Average Margin
- Improving Adversarial Robustness via Guided Complement Entropy
- Detecting Adversarial Examples via Neural Fingerprinting
- Better the Devil you Know: An Analysis of Evasion Attacks using Out-of-Distribution Adversarial Examples
- Certified Robustness of Community Detection against Adversarial Structural Perturbation via Randomized Smoothing
- Truth or Backpropaganda? An Empirical Investigation of Deep Learning Theory
- Lipschitz Networks and Distributional Robustness
- HoneyModels: Machine Learning Honeypots
- Refactoring Neural Networks for Verification
- Adversarial Robustness of Supervised Sparse Coding
- Semantics Preserving Adversarial Learning
- On Training Robust PDF Malware Classifiers
- On the Need for Topology-Aware Generative Models for Manifold-Based Defenses
- Boosting the Certified Robustness of L-infinity Distance Nets
- Fast Certified Robust Training with Short Warmup
- ART: Abstraction Refinement-Guided Training for Provably Correct Neural Networks
- Improving the Tightness of Convex Relaxation Bounds for Training Certifiably Robust Classifiers
- Evading Adversarial Example Detection Defenses with Orthogonal Projected Gradient Descent
- Robustness Certificates Against Adversarial Examples for ReLU Networks
- Certified Adversarial Defenses Meet Out-of-Distribution Corruptions: Benchmarking Robustness and Simple Baselines
- Lipschitz Bounds and Provably Robust Training by Laplacian Smoothing
- A Comprehensive Evaluation Framework for Deep Model Robustness
- Rearchitecting Classification Frameworks For Increased Robustness
- Enhancing Certified Robustness via Smoothed Weighted Ensembling
- Almost Tight L0-norm Certified Robustness of Top-k Predictions against Adversarial Perturbations
- On the Sample Complexity of Adversarial Multi-Source PAC Learning
- A Method for Computing Class-wise Universal Adversarial Perturbations
- Detection as Regression: Certified Object Detection by Median Smoothing
- Improving Adversarial Robustness via Unlabeled Out-of-Domain Data
- Model-Based Robust Deep Learning: Generalizing to Natural, Out-of-Distribution Data
- Certified Robustness of Graph Neural Networks against Adversarial Structural Perturbation
- DeepSearch: A Simple and Effective Blackbox Attack for Deep Neural Networks
- CROP: Certifying Robust Policies for Reinforcement Learning through Functional Smoothing
- Robust Adversarial Learning via Sparsifying Front Ends
- Towards Understanding Limitations of Pixel Discretization Against Adversarial Attacks
- Double Backpropagation for Training Autoencoders against Adversarial Attack
- Adversarially Robust Learning for Security-Constrained Optimal Power Flow
- A Singular Value Perspective on Model Robustness
- On the Tightness of Semidefinite Relaxations for Certifying Robustness to Adversarial Examples
- RecurJac: An Efficient Recursive Algorithm for Bounding Jacobian Matrix of Neural Networks and Its Applications
- Provable robustness against all adversarial -perturbations for
- Towards Certifying L-infinity Robustness using Neural Networks with L-inf-dist Neurons
- Architecture Selection via the Trade-off Between Accuracy and Robustness
- Understanding Catastrophic Overfitting in Adversarial Training
- Adversarial Examples for Cost-Sensitive Classifiers
- DetectorGuard: Provably Securing Object Detectors against Localized Patch Hiding Attacks
- Accelerating Robustness Verification of Deep Neural Networks Guided by Target Labels
- Universal Approximation with Certified Networks
- Relaxing Local Robustness
- D-square-B: Deep Distribution Bound for Natural-looking Adversarial Attack
- Insta-RS: Instance-wise Randomized Smoothing for Improved Robustness and Accuracy
- Robust Classification using Robust Feature Augmentation
- Adversarial Robustness Guarantees for Random Deep Neural Networks
- A Robust Classification-autoencoder to Defend Outliers and Adversaries
- Certified Defense via Latent Space Randomized Smoothing with Orthogonal Encoders
- AdvFoolGen: Creating Persistent Troubles for Deep Classifiers
- When the Guard failed the Droid: A case study of Android malware
- Domain Invariant Adversarial Learning
- Robust Generalization of Quadratic Neural Networks via Function Identification
- Empirically Measuring Concentration: Fundamental Limits on Intrinsic Robustness
- 10 Security and Privacy Problems in Large Foundation Models
- Adversarial Learning with Cost-Sensitive Classes
- Resilience from Diversity: Population-based approach to harden models against adversarial attacks
- Sample Complexity of Adversarially Robust Linear Classification on Separated Data
- When are Non-Parametric Methods Robust?
- Probabilistic Verification of Fairness Properties via Concentration
- Adversarial Robustness on Image Classification with -means
- Meta Gradient Adversarial Attack
- Overestimation learning with guarantees
- Generating Structured Adversarial Attacks Using Frank-Wolfe Method
- A Multiclass Boosting Framework for Achieving Fast and Provable Adversarial Robustness
- Safety-Aware Hardening of 3D Object Detection Neural Network Systems
- Towards Evaluating and Training Verifiably Robust Neural Networks
- Neural Belief Reasoner
- Self-Gradient Networks
- Training Machine Learning Models by Regularizing their Explanations
- Adversarial attacks on neural networks through canonical Riemannian foliations
- Hidden Cost of Randomized Smoothing
- A Neuro-Inspired Autoencoding Defense Against Adversarial Perturbations
- Intriguing Properties of Input-dependent Randomized Smoothing
- Defenses Against Multi-Sticker Physical Domain Attacks on Classifiers
- Expected Tight Bounds for Robust Training
- Training Provably Robust Models by Polyhedral Envelope Regularization
- Is the Rush to Machine Learning Jeopardizing Safety? Results of a Survey
- Towards Robustness against Unsuspicious Adversarial Examples
- Improving Resistance to Adversarial Deformations by Regularizing Gradients
- Adaptive Perturbation for Adversarial Attack
- Consistent Non-Parametric Methods for Maximizing Robustness
- A Primer on Multi-Neuron Relaxation-based Adversarial Robustness Certification
- Robust Machine Learning via Privacy/Rate-Distortion Theory
- Group-Structured Adversarial Training
- Using an ensemble color space model to tackle adversarial examples
- Reachability Analysis of Neural Networks with Uncertain Parameters
- Local Lipschitz Constant Computation of ReLU-FNNs: Upper Bound Computation with Exactness Verification
- A New Family of Neural Networks Provably Resistant to Adversarial Attacks