A Simple Explanation for the Existence of Adversarial Examples with Small Hamming Distance
arXiv:1901.10861
Abstract
The existence of adversarial examples in which an imperceptible change in the input can fool well trained neural networks was experimentally discovered by Szegedy et al in 2013, who called them "Intriguing properties of neural networks". Since then, this topic had become one of the hottest research areas within machine learning, but the ease with which we can switch between any two decisions in targeted attacks is still far from being understood, and in particular it is not clear which parameters determine the number of input coordinates we have to change in order to mislead the network. In this paper we develop a simple mathematical framework which enables us to think about this baffling phenomenon from a fresh perspective, turning it into a natural consequence of the geometry of with the (Hamming) metric, which can be quantitatively analyzed. In particular, we explain why we should expect to find targeted adversarial examples with Hamming distance of roughly in arbitrarily deep neural networks which are designed to distinguish between input classes.
References in corpus (2)
Cited by in corpus (11)
- There are No Bit Parts for Sign Bits in Black-Box Attacks
- Transferable Universal Adversarial Perturbations Using Generative Models
- Thief, Beware of What Get You There: Towards Understanding Model Extraction Attack
- Most ReLU Networks Suffer from Adversarial Perturbations
- Adversarial Examples in Multi-Layer Random ReLU Networks
- Towards Identifying and closing Gaps in Assurance of autonomous Road vehicleS -- a collection of Technical Notes Part 1
- ROBY: Evaluating the Robustness of a Deep Model by its Decision Boundaries
- Adversarial robustness via stochastic regularization of neural activation sensitivity
- A single gradient step finds adversarial examples on random two-layers neural networks
- CodNN -- Robust Neural Networks From Coded Classification
- Understanding Robustness in Teacher-Student Setting: A New Perspective