Formal Guarantees on the Robustness of a Classifier against Adversarial Manipulation
arXiv:1705.08475
Abstract
Recent work has shown that state-of-the-art classifiers are quite brittle, in the sense that a small adversarial change of an originally with high confidence correctly classified input leads to a wrong classification again with high confidence. This raises concerns that such classifiers are vulnerable to attacks and calls into question their usage in safety-critical systems. We show in this paper for the first time formal guarantees on the robustness of a classifier by giving instance-specific lower bounds on the norm of the input manipulation required to change the classifier decision. Based on this analysis we propose the Cross-Lipschitz regularization functional. We show that using this form of regularization in kernel methods resp. neural networks improves the robustness of the classifier without any loss in prediction performance.
final version accepted at NIPS 2017, fixed bug in implementation of Cross-Lipschitz regularization and lower bound computation, now results are better
Cited by in corpus (83)
- Certified Defenses against Adversarial Examples
- Efficient Neural Network Robustness Certification with General Activation Functions
- Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach
- Towards Fast Computation of Certified Robustness for ReLU Networks
- Training robust neural networks using Lipschitz bounds
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary Attack
- Rethinking Softmax Cross-Entropy Loss for Adversarial Robustness
- Automatic Perturbation Analysis for Scalable Certified Robustness and Beyond
- MMA Training: Direct Input Space Margin Maximization through Adversarial Training
- Analyzing the Robustness of Nearest Neighbors to Adversarial Examples
- How to Certify Machine Learning Based Safety-critical Systems? A Systematic Literature Review
- Attacks Which Do Not Kill Training Make Adversarial Learning Stronger
- A Closer Look at Accuracy vs. Robustness
- Sensitivity Analysis of Deep Neural Networks
- Partial success in closing the gap between human and machine vision
- Controlling Neural Level Sets
- Security and Privacy Issues in Deep Learning
- A Survey of Safety and Trustworthiness of Deep Neural Networks: Verification, Testing, Adversarial Attack and Defence, and Interpretability
- Virtual Mixup Training for Unsupervised Domain Adaptation
- Understanding and Improving Fast Adversarial Training
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- Branch and Bound for Piecewise Linear Neural Network Verification
- Towards Stable and Efficient Training of Verifiably Robust Neural Networks
- Provably Robust Boosted Decision Stumps and Trees against Adversarial Attacks
- Do Wider Neural Networks Really Help Adversarial Robustness?
- Robustness Verification for Transformers
- L2-Nonexpansive Neural Networks
- Globally-Robust Neural Networks
- Software Testing for Machine Learning
- Increasing the Confidence of Deep Neural Networks by Coverage Analysis
- Divide, Denoise, and Defend against Adversarial Attacks
- Learning Security Classifiers with Verified Global Robustness Properties
- Lipschitz regularized Deep Neural Networks generalize and are adversarially robust
- High Dimensional Spaces, Deep Learning and Adversarial Examples
- Improved robustness to adversarial examples using Lipschitz regularization of the loss
- There is Limited Correlation between Coverage and Robustness for Deep Neural Networks
- Evading classifiers in discrete domains with provable optimality guarantees
- Provable Certificates for Adversarial Examples: Fitting a Ball in the Union of Polytopes
- Adversarial Framework with Certified Robustness for Time-Series Domain via Statistical Features
- Learning Lyapunov Functions for Piecewise Affine Systems with Neural Network Controllers
- Achieving Adversarial Robustness via Sparsity
- Bag of Tricks for Adversarial Training
- Defence against adversarial attacks using classical and quantum-enhanced Boltzmann machines
- Manifold Regularization for Locally Stable Deep Neural Networks
- Understanding Generalization in Adversarial Training via the Bias-Variance Decomposition
- Calibrated Surrogate Losses for Adversarially Robust Classification
- Efficient Adversarial Training with Transferable Adversarial Examples
- On Mixup Regularization
- Understanding Adversarial Robustness: The Trade-off between Minimum and Average Margin
- Efficient Robustness Certificates for Discrete Data: Sparsity-Aware Randomized Smoothing for Graphs, Images and More
- Fantastic Four: Differentiable Bounds on Singular Values of Convolution Layers
- Investigating Decision Boundaries of Trained Neural Networks
- Is Robustness the Cost of Accuracy? -- A Comprehensive Study on the Robustness of 18 Deep Image Classification Models
- RNAS-CL: Robust Neural Architecture Search by Cross-Layer Knowledge Distillation
- Rearchitecting Classification Frameworks For Increased Robustness
- Robustness Certificates Against Adversarial Examples for ReLU Networks
- Robust Adversarial Learning via Sparsifying Front Ends
- CROP: Certifying Robust Policies for Reinforcement Learning through Functional Smoothing
- RecurJac: An Efficient Recursive Algorithm for Bounding Jacobian Matrix of Neural Networks and Its Applications
- Provable robustness against all adversarial -perturbations for
- A Survey on Trust Metrics for Autonomous Robotic Systems
- A Game Theoretic Analysis of Additive Adversarial Attacks and Defenses
- Bridging the Gap Between Adversarial Robustness and Optimization Bias
- Learning Diverse-Structured Networks for Adversarial Robustness
- An Empirical Study on the Relation between Network Interpretability and Adversarial Robustness
- Certifiably Robust Variational Autoencoders
- A Unified Game-Theoretic Interpretation of Adversarial Robustness
- A Robust Classification-autoencoder to Defend Outliers and Adversaries
- What it Thinks is Important is Important: Robustness Transfers through Input Gradients
- Robustness Guarantees for Deep Neural Networks on Videos
- Recovery Guarantees for Compressible Signals with Adversarial Noise
- SPADE: A Spectral Method for Black-Box Adversarial Robustness Evaluation
- Variational Autoencoders: A Harmonic Perspective
- Resilience from Diversity: Population-based approach to harden models against adversarial attacks
- Robust and Information-theoretically Safe Bias Classifier against Adversarial Attacks
- A general framework for defining and optimizing robustness
- Hidden Cost of Randomized Smoothing
- Intriguing Properties of Input-dependent Randomized Smoothing
- Noisy Feature Mixup
- Bridged Adversarial Training
- Recent Advances in Large Margin Learning
- Tightening the Approximation Error of Adversarial Risk with Auto Loss Function Search
- Consistent Non-Parametric Methods for Maximizing Robustness