Training Provably Robust Models by Polyhedral Envelope Regularization
arXiv:1912.04792
Abstract
Training certifiable neural networks enables one to obtain models with robustness guarantees against adversarial attacks. In this work, we introduce a framework to bound the adversary-free region in the neighborhood of the input data by a polyhedral envelope, which yields finer-grained certified robustness. We further introduce polyhedral envelope regularization (PER) to encourage larger polyhedral envelopes and thus improve the provable robustness of the models. We demonstrate the flexibility and effectiveness of our framework on standard benchmarks; it applies to networks of different architectures and general activation functions. Compared with the state-of-the-art methods, PER has very little computational overhead and better robustness guarantees without over-regularizing the model.
References in corpus (17)
- Certified Adversarial Robustness via Randomized Smoothing
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
- Using Pre-Training Can Improve Model Robustness and Uncertainty
- Efficient Neural Network Robustness Certification with General Activation Functions
- On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models
- Provably Robust Deep Learning via Adversarially Trained Smoothed Classifiers
- Improving Adversarial Robustness via Promoting Ensemble Diversity
- Semidefinite relaxations for certifying robustness to adversarial examples
- Are Labels Required for Improving Adversarial Robustness?
- Efficient Formal Safety Analysis of Neural Networks
- Uncovering the Limits of Adversarial Training against Norm-Bounded Adversarial Examples
- Training for Faster Adversarial Robustness Verification via Inducing ReLU Stability
- Rethinking Softmax Cross-Entropy Loss for Adversarial Robustness
- Provable Robustness of ReLU networks via Maximization of Linear Regions
- Metric Learning for Adversarial Robustness
- On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome Them
- Harnessing the Vulnerability of Latent Layers in Adversarially Trained Models