Theoretically Principled Trade-off between Robustness and Accuracy
arXiv:1901.08573
Abstract
We identify a trade-off between robustness and accuracy that serves as a guiding principle in the design of defenses against adversarial examples. Although this problem has been widely studied empirically, much remains unknown concerning the theory underlying this trade-off. In this work, we decompose the prediction error for adversarial examples (robust error) as the sum of the natural (classification) error and boundary error, and provide a differentiable upper bound using the theory of classification-calibrated loss, which is shown to be the tightest possible upper bound uniform over all probability distributions and measurable predictors. Inspired by our theoretical analysis, we also design a new defense method, TRADES, to trade adversarial robustness off against accuracy. Our proposed algorithm performs well experimentally in real-world datasets. The methodology is the foundation of our entry to the NeurIPS 2018 Adversarial Vision Challenge in which we won the 1st place out of ~2,000 submissions, surpassing the runner-up approach by in terms of mean perturbation distance.
Appeared in ICML 2019; the winning methodology of the NeurIPS 2018 Adversarial Vision Challenge
References in corpus (2)
Cited by in corpus (100)
- DeepRobust: A PyTorch Library for Adversarial Attacks and Defenses
- Towards Evaluating the Robustness of Deep Diagnostic Models by Adversarial Attack
- An Alternative Surrogate Loss for PGD-based Adversarial Testing
- Learning from others' mistakes: Avoiding dataset biases without modeling them
- Unlearnable Examples: Making Personal Data Unexploitable
- Robust Reinforcement Learning on State Observations with Learned Optimal Adversary
- Overfitting in adversarially robust deep learning
- Triple Wins: Boosting Accuracy, Robustness and Efficiency Together by Enabling Input-Adaptive Inference
- MaxUp: A Simple Way to Improve Generalization of Neural Network Training
- Understanding Real-world Threats to Deep Learning Models in Android Apps
- How Should Pre-Trained Language Models Be Fine-Tuned Towards Adversarial Robustness?
- A Closer Look at the Robustness of Vision-and-Language Pre-trained Models
- Bridging Adversarial Robustness and Gradient Interpretability
- Towards Understanding Fast Adversarial Training
- Probabilistic Deep Learning to Quantify Uncertainty in Air Quality Forecasting
- Recent advances in deep learning theory
- Towards Compact and Robust Deep Neural Networks
- Dual Manifold Adversarial Robustness: Defense against Lp and non-Lp Adversarial Attacks
- What You See is Not What the Network Infers: Detecting Adversarial Examples Based on Semantic Contradiction
- A Simple Fine-tuning Is All You Need: Towards Robust Deep Learning Via Adversarial Fine-tuning
- Towards Understanding the Adversarial Vulnerability of Skeleton-based Action Recognition
- Random Hypervolume Scalarizations for Provable Multi-Objective Black Box Optimization
- Two Souls in an Adversarial Image: Towards Universal Adversarial Example Detection using Multi-view Inconsistency
- EllSeg-Gen, towards Domain Generalization for head-mounted eyetracking
- Multitask Learning Strengthens Adversarial Robustness
- Neural Architecture Dilation for Adversarial Robustness
- Adversarial Feature Augmentation and Normalization for Visual Recognition
- Jacobian Adversarially Regularized Networks for Robustness
- Fighting Gradients with Gradients: Dynamic Defenses against Adversarial Attacks
- Attribution-driven Causal Analysis for Detection of Adversarial Examples
- DAmageNet: A Universal Adversarial Dataset
- Second-Order Provable Defenses against Adversarial Attacks
- Sharp Statistical Guarantees for Adversarially Robust Gaussian Classification
- Guided Interpolation for Adversarial Training
- Regularizers for Single-step Adversarial Training
- Adversarial robustness via robust low rank representations
- Input Hessian Regularization of Neural Networks
- AccelAT: A Framework for Accelerating the Adversarial Training of Deep Neural Networks through Accuracy Gradient
- Are Interpretations Fairly Evaluated? A Definition Driven Pipeline for Post-Hoc Interpretability
- LAFEAT: Piercing Through Adversarial Defenses with Latent Features
- Adaptive Feature Alignment for Adversarial Training
- Analysis and Applications of Class-wise Robustness in Adversarial Training
- A Singular Value Perspective on Model Robustness
- Benchmarking adversarial attacks and defenses for time-series data
- Improved Adversarial Training via Learned Optimizer
- Robustness Out of the Box: Compositional Representations Naturally Defend Against Black-Box Patch Attacks
- A Frequency Perspective of Adversarial Robustness
- On Norm-Agnostic Robustness of Adversarial Training
- Adversarial and Natural Perturbations for General Robustness
- Defending against substitute model black box adversarial attacks with the 01 loss
- Vulnerability Under Adversarial Machine Learning: Bias or Variance?
- Contemplating real-world object classification
- Natural Perturbed Training for General Robustness of Neural Network Classifiers
- Understanding the Error in Evaluating Adversarial Robustness
- On the human-recognizability phenomenon of adversarially trained deep image classifiers
- Smooth Imitation Learning via Smooth Costs and Smooth Policies
- Residual Error: a New Performance Measure for Adversarial Robustness
- Towards the Memorization Effect of Neural Networks in Adversarial Training
- Robusta: Robust AutoML for Feature Selection via Reinforcement Learning
- Augmenting Model Robustness with Transformation-Invariant Attacks
- Identification of Attack-Specific Signatures in Adversarial Examples
- On -norm Robustness of Ensemble Stumps and Trees
- Robust binary classification with the 01 loss
- Building Compact and Robust Deep Neural Networks with Toeplitz Matrices
- Robustness from Simple Classifiers
- Cross-domain Cross-architecture Black-box Attacks on Fine-tuned Models with Transferred Evolutionary Strategies
- A Multiclass Boosting Framework for Achieving Fast and Provable Adversarial Robustness
- On the transferability of adversarial examples between convex and 01 loss models
- The Effects of Image Distribution and Task on Adversarial Robustness
- Self-Gradient Networks
- Extreme Value Preserving Networks
- Understanding Adversarial Behavior of DNNs by Disentangling Non-Robust and Robust Components in Performance Metric
- Graph Interpolating Activation Improves Both Natural and Robust Accuracies in Data-Efficient Deep Learning
- Towards A Conceptually Simple Defensive Approach for Few-shot classifiers Against Adversarial Support Samples
- Bridged Adversarial Training
- Adversarial Boot Camp: label free certified robustness in one epoch
- Neural Network Repair with Reachability Analysis
- Identifying Layers Susceptible to Adversarial Attacks
- Adversarial Training: embedding adversarial perturbations into the parameter space of a neural network to build a robust system
- Less is More: Feature Selection for Adversarial Robustness with Compressive Counter-Adversarial Attacks
- Adaptive versus Standard Descent Methods and Robustness Against Adversarial Examples
- Robust Regularization with Adversarial Labelling of Perturbed Samples
- Balancing Robustness and Sensitivity using Feature Contrastive Learning
- One Man's Trash is Another Man's Treasure: Resisting Adversarial Examples by Adversarial Examples
- Feature Losses for Adversarial Robustness
- Unique properties of adversarially trained linear classifiers on Gaussian data
- Sparse Coding Frontend for Robust Neural Networks
- Class-Aware Domain Adaptation for Improving Adversarial Robustness
- Regularized Training and Tight Certification for Randomized Smoothed Classifier with Provable Robustness
- Robust Face Verification via Disentangled Representations
- Improving the Certified Robustness of Neural Networks via Consistency Regularization
- Training Efficiency and Robustness in Deep Learning
- Decoder-free Robustness Disentanglement without (Additional) Supervision
- A Neuro-Inspired Autoencoding Defense Against Adversarial Perturbations
- Towards Natural Robustness Against Adversarial Examples
- Towards adversarial robustness with 01 loss neural networks
- Likelihood Landscapes: A Unifying Principle Behind Many Adversarial Defenses
- Adversarial Training with Stochastic Weight Average
- Semantics-Preserving Adversarial Training
- Adversarial Robustness Across Representation Spaces