Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach
arXiv:1801.10578
Abstract
The robustness of neural networks to adversarial examples has received great attention due to security implications. Despite various attack approaches to crafting visually imperceptible adversarial examples, little has been developed towards a comprehensive measure of robustness. In this paper, we provide a theoretical justification for converting robustness analysis into a local Lipschitz constant estimation problem, and propose to use the Extreme Value Theory for efficient evaluation. Our analysis yields a novel robustness metric called CLEVER, which is short for Cross Lipschitz Extreme Value for nEtwork Robustness. The proposed CLEVER score is attack-agnostic and computationally feasible for large neural networks. Experimental results on various networks, including ResNet, Inception-v3 and MobileNet, show that (i) CLEVER is aligned with the robustness indication measured by the and norms of adversarial examples from powerful attacks, and (ii) defended networks using defensive distillation or bounded ReLU indeed achieve better CLEVER scores. To the best of our knowledge, CLEVER is the first attack-independent robustness metric that can be applied to any neural network classifier.
Accepted by Sixth International Conference on Learning Representations (ICLR 2018). Tsui-Wei Weng and Huan Zhang contributed equally
References in corpus (4)
Cited by in corpus (23)
- advertorch v0.1: An Adversarial Robustness Toolbox based on PyTorch
- Opportunities and Challenges in Deep Learning Adversarial Robustness: A Survey
- The RFML Ecosystem: A Look at the Unique Challenges of Applying Deep Learning to Radio Frequency Applications
- There is Limited Correlation between Coverage and Robustness for Deep Neural Networks
- Discovering and Explaining the Representation Bottleneck of DNNs
- A Formalization of Robustness for Deep Neural Networks
- Orthogonalizing Convolutional Layers with the Cayley Transform
- Neural Architecture Dilation for Adversarial Robustness
- Robustness Certificates Against Adversarial Examples for ReLU Networks
- Markov-Lipschitz Deep Learning
- Data Quality Matters For Adversarial Training: An Empirical Study
- Safe Predictors for Enforcing Input-Output Specifications
- Uncertainty Quantification and Confidence Intervals for Naive Rare-Event Estimators
- Abstraction and Symbolic Execution of Deep Neural Networks with Bayesian Approximation of Hidden Features
- A Survey on Trust Metrics for Autonomous Robotic Systems
- A Geometrical Approach to Evaluate the Adversarial Robustness of Deep Neural Networks
- Deterministic Gaussian Averaged Neural Networks
- Mitigating Deep Learning Vulnerabilities from Adversarial Examples Attack in the Cybersecurity Domain
- Resilience from Diversity: Population-based approach to harden models against adversarial attacks
- Tightening the Approximation Error of Adversarial Risk with Auto Loss Function Search
- PipeSim: Trace-driven Simulation of Large-Scale AI Operations Platforms
- Adversarial Machine Learning for Cybersecurity and Computer Vision: Current Developments and Challenges
- An Efficient and Margin-Approaching Zero-Confidence Adversarial Attack