Locally Enhanced Self-Attention: Combining Self-Attention and Convolution as Local and Context Terms
arXiv:2107.05637
Abstract
Self-Attention has become prevalent in computer vision models. Inspired by fully connected Conditional Random Fields (CRFs), we decompose self-attention into local and context terms. They correspond to the unary and binary terms in CRF and are implemented by attention mechanisms with projection matrices. We observe that the unary terms only make small contributions to the outputs, and meanwhile standard CNNs that rely solely on the unary terms achieve great performances on a variety of tasks. Therefore, we propose Locally Enhanced Self-Attention (LESA), which enhances the unary term by incorporating it with convolutions, and utilizes a fusion module to dynamically couple the unary and binary operations. In our experiments, we replace the self-attention modules with LESA. The results on ImageNet and COCO show the superiority of LESA over convolution and self-attention baselines for the tasks of image recognition, object detection, and instance segmentation. The code is made publicly available.
References in corpus (12)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- PVT v2: Improved Baselines with Pyramid Vision Transformer
- MMDetection: Open MMLab Detection Toolbox and Benchmark
- Conditional Positional Encodings for Vision Transformers
- LocalViT: Analyzing Locality in Vision Transformers
- CvT: Introducing Convolutions to Vision Transformers
- Lite Transformer with Long-Short Range Attention
- Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth
- Scaling Wide Residual Networks for Panoptic Segmentation