Certified Adversarial Robustness via Randomized Smoothing
arXiv:1902.02918
Abstract
We show how to turn any classifier that classifies well under Gaussian noise into a new classifier that is certifiably robust to adversarial perturbations under the norm. This "randomized smoothing" technique has been proposed recently in the literature, but existing guarantees are loose. We prove a tight robustness guarantee in norm for smoothing with Gaussian noise. We use randomized smoothing to obtain an ImageNet classifier with e.g. a certified top-1 accuracy of 49% under adversarial perturbations with norm less than 0.5 (=127/255). No certified defense has been shown feasible on ImageNet except for smoothing. On smaller-scale datasets where competing approaches to certified robustness are viable, smoothing delivers higher certified accuracies. Our strong empirical results suggest that randomized smoothing is a promising direction for future research into adversarially robust classification. Code and models are available at http://github.com/locuslab/smoothing.
ICML 2019
Cited by in corpus (104)
- Fast is better than free: Revisiting adversarial training
- Invisible for both Camera and LiDAR: Security of Multi-Sensor Fusion based Perception in Autonomous Driving Under Physical-World Attacks
- Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead
- Measuring Robustness to Natural Distribution Shifts in Image Classification
- Are Perceptually-Aligned Gradients a General Property of Robust Classifiers?
- Theoretical evidence for adversarial robustness through randomization
- Triple Wins: Boosting Accuracy, Robustness and Efficiency Together by Enabling Input-Adaptive Inference
- Adversarial purification with Score-based generative models
- Certified Robustness for Top-k Predictions against Adversarial Perturbations via Randomized Smoothing
- Defense against Adversarial Attacks in NLP via Dirichlet Neighborhood Ensemble
- Boosting Adversarial Transferability through Enhanced Momentum
- On the Resilience of Biometric Authentication Systems against Random Inputs
- Towards Frequency-Based Explanation for Robust CNN
- A Survey of Black-Box Adversarial Attacks on Computer Vision Models
- Efficient Certified Defenses Against Patch Attacks on Image Classifiers
- Dual Manifold Adversarial Robustness: Defense against Lp and non-Lp Adversarial Attacks
- Yet another but more efficient black-box adversarial attack: tiling and evolution strategies
- Practical Detection of Trojan Neural Networks: Data-Limited and Data-Free Cases
- If a Human Can See It, So Should Your System: Reliability Requirements for Machine Vision Components
- Exploiting Excessive Invariance caused by Norm-Bounded Adversarial Robustness
- Adversarial training in communication constrained federated learning
- Random Smoothing Might be Unable to Certify Robustness for High-Dimensional Images
- VSMask: Defending Against Voice Synthesis Attack via Real-Time Predictive Perturbation
- ScaleCert: Scalable Certified Defense against Adversarial Patches with Sparse Superficial Layers
- Two Souls in an Adversarial Image: Towards Universal Adversarial Example Detection using Multi-view Inconsistency
- Provably Efficient Black-Box Action Poisoning Attacks Against Reinforcement Learning
- Attacking Adversarial Attacks as A Defense
- A Research Agenda: Dynamic Models to Defend Against Correlated Attacks
- Orthogonalizing Convolutional Layers with the Cayley Transform
- Adversarial Machine Learning for 5G Communications Security
- SoK: Machine Learning Governance
- Adversarial Token Attacks on Vision Transformers
- Adversarial Feature Augmentation and Normalization for Visual Recognition
- Adversarial Visual Robustness by Causal Intervention
- Fighting Gradients with Gradients: Dynamic Defenses against Adversarial Attacks
- Robust Encodings: A Framework for Combating Adversarial Typos
- PRECAD: Privacy-Preserving and Robust Federated Learning via Crypto-Aided Differential Privacy
- Mask-GVAE: Blind Denoising Graphs via Partition
- Transferred Discrepancy: Quantifying the Difference Between Representations
- Detection Defense Against Adversarial Attacks with Saliency Map
- Interpretable Graph Capsule Networks for Object Recognition
- Quantifying and Improving Transferability in Domain Generalization
- Local Reweighting for Adversarial Training
- -ML: Mitigating Adversarial Examples via Ensembles of Topologically Manipulated Classifiers
- Adversarial robustness via robust low rank representations
- Improved Autoregressive Modeling with Distribution Smoothing
- Understanding Adversarial Examples from the Mutual Influence of Images and Perturbations
- Universal Adversarial Perturbations Through the Lens of Deep Steganography: Towards A Fourier Perspective
- Training Certifiably Robust Neural Networks with Efficient Local Lipschitz Bounds
- Adversarial Training Helps Transfer Learning via Better Representations
- Random Transformation of Image Brightness for Adversarial Attack
- A unified view on differential privacy and robustness to adversarial examples
- Where is the Bottleneck of Adversarial Learning with Unlabeled Data?
- Analyzing Accuracy Loss in Randomized Smoothing Defenses
- Connecting Interpretability and Robustness in Decision Trees through Separation
- Certifying Joint Adversarial Robustness for Model Ensembles
- Auditing AI models for Verified Deployment under Semantic Specifications
- Mixed Nash Equilibria in the Adversarial Examples Game
- Adversarially Robust Learning with Unknown Perturbation Sets
- RoMA: Robust Model Adaptation for Offline Model-based Optimization
- Provable Robust Classification via Learned Smoothed Densities
- Understanding and Diagnosing Vulnerability under Adversarial Attacks
- Multi-stage Optimization based Adversarial Training
- Towards Adversarial Robustness via Transductive Learning
- Relaxing Local Robustness
- Perceptually Constrained Adversarial Attacks
- Immuno-mimetic Deep Neural Networks (Immuno-Net)
- A Multiclass Boosting Framework for Achieving Fast and Provable Adversarial Robustness
- Provable Robustness of Adversarial Training for Learning Halfspaces with Noise
- Network Moments: Extensions and Sparse-Smooth Attacks
- Building Compact and Robust Deep Neural Networks with Toeplitz Matrices
- HASI: Hardware-Accelerated Stochastic Inference, A Defense Against Adversarial Machine Learning Attacks
- Adversarial Purification through Representation Disentanglement
- Adversarial Attacks on ML Defense Models Competition
- Practical Convex Formulation of Robust One-hidden-layer Neural Network Training
- When Should You Defend Your Classifier -- A Game-theoretical Analysis of Countermeasures against Adversarial Examples
- Identifying and Exploiting Structures for Reliable Deep Learning
- Deep Adversarially-Enhanced k-Nearest Neighbors
- Leave-one-out Unfairness
- Countering Adversarial Examples: Combining Input Transformation and Noisy Training
- Hard-label Manifolds: Unexpected Advantages of Query Efficiency for Finding On-manifold Adversarial Examples
- Privacy Preserving Recalibration under Domain Shift
- Federated Learning without Revealing the Decision Boundaries
- Constructing a provably adversarially-robust classifier from a high accuracy one
- Certified Robustness of Graph Classification against Topology Attack with Randomized Smoothing
- Robustifying Adversarial Training to the Union of Perturbation Models
- Adversarial Training with Stochastic Weight Average
- Semantics-Preserving Adversarial Training
- Improved Estimation of Concentration Under -Norm Distance Metrics Using Half Spaces
- Adaptive Verifiable Training Using Pairwise Class Similarity
- Learning to Separate Clusters of Adversarial Representations for Robust Adversarial Detection
- Stochastic-HMDs: Adversarial Resilient Hardware Malware Detectors through Voltage Over-scaling
- Deterministic Certification to Adversarial Attacks via Bernstein Polynomial Approximation
- Extreme Value Preserving Networks
- Simpler Certified Radius Maximization by Propagating Covariances
- Defense-friendly Images in Adversarial Attacks: Dataset and Metrics for Perturbation Difficulty
- Training Efficiency and Robustness in Deep Learning
- Adversarial Machine Learning in Text Analysis and Generation
- ε-weakened Robustness of Deep Neural Networks
- Adversarial Boot Camp: label free certified robustness in one epoch
- Improving the Certified Robustness of Neural Networks via Consistency Regularization
- Calibrated Adversarial Training
- Towards Assessment of Randomized Smoothing Mechanisms for Certifying Adversarial Robustness
- Improving Adversarial Robustness for Free with Snapshot Ensemble