LocalViT: Analyzing Locality in Vision Transformers
arXiv:2104.05707
Abstract
The aim of this paper is to study the influence of locality mechanisms in vision transformers. Transformers originated from machine translation and are particularly good at modelling long-range dependencies within a long sequence. Although the global interaction between the token embeddings could be well modelled by the self-attention mechanism of transformers, what is lacking is a locality mechanism for information exchange within a local region. In this paper, locality mechanism is systematically investigated by carefully designed controlled experiments. We add locality to vision transformers into the feed-forward network. This seemingly simple solution is inspired by the comparison between feed-forward networks and inverted residual blocks. The importance of locality mechanisms is validated in two ways: 1) A wide range of design choices (activation function, layer placement, expansion ratio) are available for incorporating locality mechanisms and proper choices can lead to a performance gain over the baseline, and 2) The same locality mechanism is successfully applied to vision transformers with different architecture designs, which shows the generalization of the locality concept. For ImageNet2012 classification, the locality-enhanced transformers outperform the baselines Swin-T, DeiT-T, and PVT-T by 1.0%, 2.6% and 3.1% with a negligible increase in the number of parameters and computational effort. Code is available at https://github.com/ofsoundof/LocalViT.
References in corpus (18)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
- Deep Residual Learning for Image Recognition
- MobileNetV2: Inverted Residuals and Linear Bottlenecks
- Densely Connected Convolutional Networks
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Transformer in Transformer
- Fully Convolutional Networks for Semantic Segmentation
- Rich feature hierarchies for accurate object detection and semantic segmentation
- TransTrack: Multiple Object Tracking with Transformer
- Non-Local Recurrent Network for Image Restoration
- Lite Transformer with Long-Short Range Attention
- Pre-Trained Image Processing Transformer
- Rethinking Attention with Performers
- DeLighT: Deep and Light-weight Transformer
Cited by in corpus (45)
- A Survey on Visual Transformer
- Transformers in Vision: A Survey
- PVT v2: Improved Baselines with Pyramid Vision Transformer
- Vision Transformers for Single Image Dehazing
- A survey of the Vision Transformers and their CNN-Transformer based Variants
- Filter-enhanced MLP is All You Need for Sequential Recommendation
- P2T: Pyramid Pooling Transformer for Scene Understanding
- Practical Blind Image Denoising via Swin-Conv-UNet and Data Synthesis
- Pneumonia Detection on chest X-ray images Using Ensemble of Deep Convolutional Neural Networks
- MISSFormer: An Effective Medical Image Segmentation Transformer
- TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding Network
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer
- HRFormer: High-Resolution Transformer for Dense Prediction
- Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer
- Uformer: A General U-Shaped Transformer for Image Restoration
- RegionViT: Regional-to-Local Attention for Vision Transformers
- Tracker Meets Night: A Transformer Enhancer for UAV Tracking
- AS-MLP: An Axial Shifted MLP Architecture for Vision
- SwinIR: Image Restoration Using Swin Transformer
- A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP
- LIT-Former: Linking In-plane and Through-plane Transformers for Simultaneous CT Image Denoising and Deblurring
- Depth-Wise Convolutions in Vision Transformers for Efficient Training on Small Datasets
- A Survey of Visual Transformers
- Refiner: Refining Self-attention for Vision Transformers
- Fully Transformer Networks for Semantic Image Segmentation
- Hierarchical Vision Transformers for Cardiac Ejection Fraction Estimation
- STB-VMM: Swin Transformer Based Video Motion Magnification
- Local-to-Global Self-Attention in Vision Transformers
- DnSwin: Toward Real-World Denoising via Continuous Wavelet Sliding-Transformer
- DuDoTrans: Dual-Domain Transformer Provides More Attention for Sinogram Restoration in Sparse-View CT Reconstruction
- HRPVT: High-Resolution Pyramid Vision Transformer for medium and small-scale human pose estimation
- Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image Recognition
- OH-Former: Omni-Relational High-Order Transformer for Person Re-Identification
- Pruning Self-attentions into Convolutional Layers in Single Path
- AdaViT: Adaptive Vision Transformers for Efficient Image Recognition
- Analogous to Evolutionary Algorithm: Designing a Unified Sequence Model
- Towards Robust Vision Transformer
- gSwin: Gated MLP Vision Model with Hierarchical Structure of Shifted Window
- More than Encoder: Introducing Transformer Decoder to Upsample
- Locally Enhanced Self-Attention: Combining Self-Attention and Convolution as Local and Context Terms
- Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation
- KVT: k-NN Attention for Boosting Vision Transformers
- Ripple Attention for Visual Perception with Sub-quadratic Complexity
- Exploring and Improving Mobile Level Vision Transformers