Certifiably Robust Variational Autoencoders
arXiv:2102.07559
Abstract
We introduce an approach for training Variational Autoencoders (VAEs) that are certifiably robust to adversarial attack. Specifically, we first derive actionable bounds on the minimal size of an input perturbation required to change a VAE's reconstruction by more than an allowed amount, with these bounds depending on certain key parameters such as the Lipschitz constants of the encoder and decoder. We then show how these parameters can be controlled, thereby providing a mechanism to ensure \textit{a priori} that a VAE will attain a desired level of robustness. Moreover, we extend this to a complete practical approach for training such VAEs to ensure our criteria are met. Critically, our method allows one to specify a desired level of robustness \emph{upfront} and then train a VAE that is guaranteed to achieve this robustness. We further demonstrate that these Lipschitz--constrained VAEs are more robust to attack than standard VAEs in practice.
12 pages and appendix
References in corpus (9)
- Certified Adversarial Robustness via Randomized Smoothing
- Provably Robust Deep Learning via Adversarially Trained Smoothed Classifiers
- A Closer Look at Accuracy vs. Robustness
- Adversarial Images for Variational Autoencoders
- The continuous Bernoulli: fixing a pervasive error in variational autoencoders
- Preventing Gradient Attenuation in Lipschitz Constrained Convolutional Networks
- Autoencoding Variational Autoencoder
- On Implicit Regularization in -VAEs
- Improving VAEs' Robustness to Adversarial Attack