Rethinking Spatial Dimensions of Vision Transformers
arXiv:2103.16302
Abstract
Vision Transformer (ViT) extends the application range of transformers from language processing to computer vision tasks as being an alternative architecture against the existing convolutional neural networks (CNN). Since the transformer-based architecture has been innovative for computer vision modeling, the design convention towards an effective architecture has been less studied yet. From the successful design principles of CNN, we investigate the role of spatial dimension conversion and its effectiveness on transformer-based architecture. We particularly attend to the dimension reduction principle of CNNs; as the depth increases, a conventional CNN increases channel dimension and decreases spatial dimensions. We empirically show that such a spatial dimension reduction is beneficial to a transformer architecture as well, and propose a novel Pooling-based Vision Transformer (PiT) upon the original ViT model. We show that PiT achieves the improved model capability and generalization performance against ViT. Throughout the extensive experiments, we further show PiT outperforms the baseline on several tasks such as image classification, object detection, and robustness evaluation. Source codes and ImageNet models are available at https://github.com/naver-ai/pit
ICCV 2021 camera-ready version
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing
- Noise or Signal: The Role of Image Backgrounds in Object Recognition
- LambdaNetworks: Modeling Long-Range Interactions Without Attention
- Rethinking Natural Adversarial Examples for Classification Models
Cited by in corpus (19)
- A survey of the Vision Transformers and their CNN-Transformer based Variants
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
- All Tokens Matter: Token Labeling for Training Better Vision Transformers
- RegionViT: Regional-to-Local Attention for Vision Transformers
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
- S-MLP: Spatial-Shift MLP Architecture for Vision
- VOLO: Vision Outlooker for Visual Recognition
- Visual Parser: Representing Part-whole Hierarchies with Transformers
- HAT: Hierarchical Aggregation Transformers for Person Re-identification
- Searching for Efficient Multi-Stage Vision Transformers
- UFO-ViT: High Performance Linear Vision Transformer without Softmax
- Towards Robust Vision Transformer
- An Image Patch is a Wave: Phase-Aware Vision MLP
- KVT: k-NN Attention for Boosting Vision Transformers
- Transformed CNNs: recasting pre-trained convolutional layers with self-attention
- Efficient Video Transformers with Spatial-Temporal Token Selection
- Measure Twice, Cut Once: Quantifying Bias and Fairness in Deep Neural Networks
- Semi-Supervised Vision Transformers
- Ripple Attention for Visual Perception with Sub-quadratic Complexity