Intriguing Properties of Vision Transformers
arXiv:2105.10497
Abstract
Vision transformers (ViT) have demonstrated impressive performance across various machine vision problems. These models are based on multi-head self-attention mechanisms that can flexibly attend to a sequence of image patches to encode contextual cues. An important question is how such flexibility in attending image-wide context conditioned on a given patch can facilitate handling nuisances in natural images e.g., severe occlusions, domain shifts, spatial permutations, adversarial and natural perturbations. We systematically study this question via an extensive set of experiments encompassing three ViT families and comparisons with a high-performing convolutional neural network (CNN). We show and analyze the following intriguing properties of ViT: (a) Transformers are highly robust to severe occlusions, perturbations and domain shifts, e.g., retain as high as 60% top-1 accuracy on ImageNet even after randomly occluding 80% of the image content. (b) The robust performance to occlusions is not due to a bias towards local textures, and ViTs are significantly less biased towards textures compared to CNNs. When properly trained to encode shape-based features, ViTs demonstrate shape recognition capability comparable to that of human visual system, previously unmatched in the literature. (c) Using ViTs to encode shape representation leads to an interesting consequence of accurate semantic segmentation without pixel-level supervision. (d) Off-the-shelf features from a single ViT model can be combined to create a feature ensemble, leading to high accuracy rates across a range of classification datasets in both traditional and few-shot learning paradigms. We show effective features of ViTs are due to flexible and dynamic receptive fields possible via the self-attention mechanism.
NeurIPS'21 (Spotlight), Code: https://git.io/Js15X
References in corpus (6)
- Distilling the Knowledge in a Neural Network
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Stand-Alone Self-Attention in Vision Models
- A Study of Face Obfuscation in ImageNet
- Are Convolutional Neural Networks or Transformers more like human vision?
- Does enhanced shape bias improve neural network robustness to common corruptions?
Cited by in corpus (35)
- Efficient Training of Audio Transformers with Patchout
- Video Transformers: A Survey
- Improving Chest X-Ray Report Generation by Leveraging Warm Starting
- CDTrans: Cross-domain Transformer for Unsupervised Domain Adaptation
- Do Vision Transformers See Like Convolutional Neural Networks?
- Localizing Objects with Self-Supervised Transformers and no Labels
- Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
- Partial success in closing the gap between human and machine vision
- Swin-transformer-yolov5 For Real-time Wine Grape Bunch Detection
- RobustART: Benchmarking Robustness on Architecture Design and Training Techniques
- On the Role of ViT and CNN in Semantic Communications: Analysis and Prototype Validation
- Towards Evaluating Explanations of Vision Transformers for Medical Imaging
- A Transformer-based Generative Adversarial Network for Brain Tumor Segmentation
- SwinCross: Cross-modal Swin Transformer for Head-and-Neck Tumor Segmentation in PET/CT Images
- Assaying Out-Of-Distribution Generalization in Transfer Learning
- Exploiting Shape Cues for Weakly Supervised Semantic Segmentation
- Vision Transformers for Small Histological Datasets Learned through Knowledge Distillation
- Exploring Advances in Transformers and CNN for Skin Lesion Diagnosis on Small Datasets
- Pruning Self-attentions into Convolutional Layers in Single Path
- Transfer Learning Gaussian Anomaly Detection by Fine-tuning Representations
- PatchCensor: Patch Robustness Certification for Transformers via Exhaustive Testing
- Understanding and Improving Robustness of Vision Transformers through Patch-based Negative Augmentation
- The Treachery of Images: Bayesian Scene Keypoints for Deep Policy Learning in Robotic Manipulation
- Discrete Representations Strengthen Vision Transformer Robustness
- A Temporal-Spectral Fusion Transformer with Subject-Specific Adapter for Enhancing RSVP-BCI Decoding
- How and When Adversarial Robustness Transfers in Knowledge Distillation?
- Pyramid Adversarial Training Improves ViT Performance
- DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview Cameras
- Efficient Video Transformers with Spatial-Temporal Token Selection
- TransMix: Attend to Mix for Vision Transformers
- Clarifying Myths About the Relationship Between Shape Bias, Accuracy, and Robustness
- To Make Yourself Invisible with Adversarial Semantic Contours
- Benchmarking the Spatial Robustness of DNNs via Natural and Adversarial Localized Corruptions
- Adversarial AutoMixup
- Improved Robustness of Vision Transformer via PreLayerNorm in Patch Embedding