Architecture Selection via the Trade-off Between Accuracy and Robustness
arXiv:1906.01354
Abstract
We provide a general framework for characterizing the trade-off between accuracy and robustness in supervised learning. We propose a method and define quantities to characterize the trade-off between accuracy and robustness for a given architecture, and provide theoretical insight into the trade-off. Specifically we introduce a simple trade-off curve, define and study an influence function that captures the sensitivity, under adversarial attack, of the optima of a given loss function. We further show how adversarial training regularizes the parameters in an over-parameterized linear model, recovering the LASSO and ridge regression as special cases, which also allows us to theoretically analyze the behavior of the trade-off curve. In experiments, we demonstrate the corresponding trade-off curves of neural networks and how they vary with respect to factors such as number of layers, neurons, and across different network structures. Such information provides a useful guideline to architecture selection.
Incorporated in a later submission. This submission is not complete in results
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning
- Understanding Black-box Predictions via Influence Functions
- Theoretically Principled Trade-off between Robustness and Accuracy
- Provable defenses against adversarial examples via the convex outer adversarial polytope
- Certified Defenses against Adversarial Examples
- Robustness May Be at Odds with Accuracy
- Certifying Some Distributional Robustness with Principled Adversarial Training
- Towards the first adversarially robust neural network model on MNIST
- On the Power of Over-parametrization in Neural Networks with Quadratic Activation