Pyramid Adversarial Training Improves ViT Performance
arXiv:2111.15121
Abstract
Aggressive data augmentation is a key component of the strong generalization capabilities of Vision Transformer (ViT). One such data augmentation technique is adversarial training (AT); however, many prior works have shown that this often results in poor clean accuracy. In this work, we present pyramid adversarial training (PyramidAT), a simple and effective technique to improve ViT's overall performance. We pair it with a "matched" Dropout and stochastic depth regularization, which adopts the same Dropout and stochastic depth configuration for the clean and adversarial samples. Similar to the improvements on CNNs by AdvProp (not directly applicable to ViT), our pyramid adversarial training breaks the trade-off between in-distribution accuracy and out-of-distribution robustness for ViT and related architectures. It leads to 1.82% absolute improvement on ImageNet clean accuracy for the ViT-B model when trained only on ImageNet-1K data, while simultaneously boosting performance on 7 ImageNet robustness metrics, by absolute numbers ranging from 1.76% to 15.68%. We set a new state-of-the-art for ImageNet-C (41.42 mCE), ImageNet-R (53.92%), and ImageNet-Sketch (41.04%) without extra data, using only the ViT-B/16 backbone and our pyramid adversarial training. Our code is publicly available at pyramidat.github.io.
Accepted to CVPR22 (oral, best paper finalist). 33 pages, including references & supplementary material
References in corpus (20)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- mixup: Beyond Empirical Risk Minimization
- Improved Regularization of Convolutional Neural Networks with Cutout
- MLP-Mixer: An all-MLP Architecture for Vision
- BEiT: BERT Pre-Training of Image Transformers
- Theoretically Principled Trade-off between Robustness and Accuracy
- ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Do ImageNet Classifiers Generalize to ImageNet?
- Intriguing Properties of Vision Transformers
- Spatially Transformed Adversarial Examples
- Adversarial Robustness through Local Linearization
- Improving Robustness Without Sacrificing Accuracy with Patch Gaussian Augmentation
- Are we done with ImageNet?
- How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
- Understanding and Mitigating the Tradeoff Between Robustness and Accuracy
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization
- AugMax: Adversarial Composition of Random Augmentations for Robust Training
- Towards Robust Vision Transformer
- On the Robustness of Vision Transformers to Adversarial Examples