Refiner: Refining Self-attention for Vision Transformers
arXiv:2106.03714
Abstract
Vision Transformers (ViTs) have shown competitive accuracy in image classification tasks compared with CNNs. Yet, they generally require much more data for model pre-training. Most of recent works thus are dedicated to designing more complex architectures or training methods to address the data-efficiency issue of ViTs. However, few of them explore improving the self-attention mechanism, a key factor distinguishing ViTs from CNNs. Different from existing works, we introduce a conceptually simple scheme, called refiner, to directly refine the self-attention maps of ViTs. Specifically, refiner explores attention expansion that projects the multi-head attention maps to a higher-dimensional space to promote their diversity. Further, refiner applies convolutions to augment local patterns of the attention maps, which we show is equivalent to a distributed local attention features are aggregated locally with learnable kernels and then globally aggregated with self-attention. Extensive experiments demonstrate that refiner works surprisingly well. Significantly, it enables ViTs to achieve 86% top-1 classification accuracy on ImageNet with only 81M parameters.
References in corpus (24)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
- Language Models are Few-Shot Learners
- MobileNetV2: Inverted Residuals and Linear Bottlenecks
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Transformer in Transformer
- Recurrent Neural Network for Text Classification with Multi-Task Learning
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- SQuAD: 100,000+ Questions for Machine Comprehension of Text
- DeepViT: Towards Deeper Vision Transformer
- LocalViT: Analyzing Locality in Vision Transformers
- High-Performance Large-Scale Image Recognition Without Normalization
- Synthesizer: Rethinking Self-Attention in Transformer Models
- Pre-Trained Image Processing Transformer
- End-to-End Object Detection with Adaptive Clustering Transformer
- Are we done with ImageNet?
- Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth
- End-to-End Video Instance Segmentation with Transformers
- Talking-Heads Attention
- Vision Transformers with Patch Diversification
- Fixing the train-test resolution discrepancy
- Learning Joint Spatial-Temporal Transformations for Video Inpainting