Adversarial Consistency and the Uniqueness of the Adversarial Bayes Classifier
arXiv:2404.17358 · doi:10.1017/S0956792525000038
Abstract
Minimizing an adversarial surrogate risk is a common technique for learning robust classifiers. Prior work showed that convex surrogate losses are not statistically consistent in the adversarial context -- or in other words, a minimizing sequence of the adversarial surrogate risk will not necessarily minimize the adversarial classification error. We connect the consistency of adversarial surrogate losses to properties of minimizers to the adversarial classification risk, known as adversarial Bayes classifiers. Specifically, under reasonable distributional assumptions, a convex surrogate loss is statistically consistent for adversarial learning iff the adversarial Bayes classifier satisfies a certain notion of uniqueness.
2 figures, 20 pages, v2: fixed typos, v3: improved organization of paper and added figures
References in corpus (12)
- Evasion Attacks against Machine Learning at Test Time
- Lower Bounds on Adversarial Robustness from Optimal Transport
- The Geometry of Adversarial Training in Binary Classification
- The Multimarginal Optimal Transport Formulation of Adversarial Multiclass Classification
- Calibration and Consistency of Adversarial Surrogate Losses
- Adversarial Training Should Be Cast as a Non-Zero-Sum Game
- Existence and Minimax Theorems for Adversarial Surrogate Risks in Binary Classification
- On the Role of Randomization in Adversarially Robust Classification
- Towards Consistency in Adversarial Classification
- Stratified Adversarial Robustness with Rejection
- On Achieving Optimal Adversarial Test Error
- Towards Calibrated Losses for Adversarial Robust Reject Option Classification