Adversarial Robustness Guarantees for Random Deep Neural Networks
arXiv:2004.05923
Abstract
The reliability of deep learning algorithms is fundamentally challenged by the existence of adversarial examples, which are incorrectly classified inputs that are extremely close to a correctly classified input. We explore the properties of adversarial examples for deep neural networks with random weights and biases, and prove that for any , the distance of any given input from the classification boundary scales as one over the square root of the dimension of the input times the norm of the input. The results are based on the recently proved equivalence between Gaussian processes and deep neural networks in the limit of infinite width of the hidden layers, and are validated with experiments on both random deep neural networks and deep neural networks trained on the MNIST and CIFAR10 datasets. The results constitute a fundamental advance in the theoretical understanding of adversarial examples, and open the way to a thorough theoretical characterization of the relation between network architecture and robustness to adversarial perturbations.
References in corpus (19)
- Deep Learning in Neural Networks: An Overview
- Delving into Transferable Adversarial Examples and Black-box Attacks
- The Loss Surfaces of Multilayer Networks
- Certified Adversarial Robustness via Randomized Smoothing
- On Exact Computation with an Infinitely Wide Neural Net
- Generalization Bounds of Stochastic Gradient Descent for Wide and Deep Neural Networks
- Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation
- A Boundary Tilting Persepective on the Phenomenon of Adversarial Examples
- Adversarial Training Can Hurt Generalization
- Adversarial Examples Are a Natural Consequence of Test Error in Noise
- Robustness of classifiers: from adversarial to random noise
- Enhanced Convolutional Neural Tangent Kernels
- Adversarial Robustness May Be at Odds With Simplicity
- A Simple Explanation for the Existence of Adversarial Examples with Small Hamming Distance
- Convergence of Adversarial Training in Overparametrized Neural Networks
- Random deep neural networks are biased towards simple functions
- Bridging Adversarial Robustness and Gradient Interpretability
- Provable Certificates for Adversarial Examples: Fitting a Ball in the Union of Polytopes
- Bayesian Adversarial Spheres: Bayesian Inference and Adversarial Examples in a Noiseless Setting