A Theoretical Framework for Robustness of (Deep) Classifiers against Adversarial Examples
arXiv:1612.00334
Abstract
Most machine learning classifiers, including deep neural networks, are vulnerable to adversarial examples. Such inputs are typically generated by adding small but purposeful modifications that lead to incorrect outputs while imperceptible to human eyes. The goal of this paper is not to introduce a single method, but to make theoretical steps towards fully understanding adversarial examples. By using concepts from topology, our theoretical analysis brings forth the key reasons why an adversarial example can fool a classifier () and adds its oracle (, like human eyes) in such analysis. By investigating the topological relationship between two (pseudo)metric spaces corresponding to predictor and oracle , we develop necessary and sufficient conditions that can determine if is always robust (strong-robust) against adversarial examples according to . Interestingly our theorems indicate that just one unnecessary feature can make not strong-robust, and the right feature representation learning is the key to getting a classifier that is both accurate and strong-robust.
38 pages , ICLR 2017 Workshop Track
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Wide Residual Networks
- Poisoning Attacks against Support Vector Machines
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Adversarial Perturbations Against Deep Neural Networks for Malware Classification
- Is feature selection secure against training data poisoning?
- Adversarial Feature Selection against Evasion Attacks
- Defensive Distillation is Not Robust to Adversarial Examples
- Crypto-Nets: Neural Networks over Encrypted Data
- Differentially Private Algorithms for Empirical Machine Learning
- Differentially- and non-differentially-private random decision trees
Cited by in corpus (8)
- Microsoft Malware Classification Challenge
- Graph Neural Networks with Continual Learning for Fake News Detection from Social Media
- Attention Meets Perturbations: Robust and Interpretable Attention with Adversarial Training
- Curls & Whey: Boosting Black-Box Adversarial Attacks
- One Bit Matters: Understanding Adversarial Examples as the Abuse of Redundancy
- Adversarial Attack Type I: Cheat Classifiers by Significant Changes
- ROBY: Evaluating the Robustness of a Deep Model by its Decision Boundaries
- Achieving Adversarial Robustness Requires An Active Teacher