A Survey of Visual Transformers
arXiv:2111.06091
Abstract
Transformer, an attention-based encoder-decoder model, has already revolutionized the field of natural language processing (NLP). Inspired by such significant achievements, some pioneering works have recently been done on employing Transformer-liked architectures in the computer vision (CV) field, which have demonstrated their effectiveness on three fundamental CV tasks (classification, detection, and segmentation) as well as multiple sensory data stream (images, point clouds, and vision-language data). Because of their competitive modeling capabilities, the visual Transformers have achieved impressive performance improvements over multiple benchmarks as compared with modern Convolution Neural Networks (CNNs). In this survey, we have reviewed over one hundred of different visual Transformers comprehensively according to three fundamental CV tasks and different data stream types, where a taxonomy is proposed to organize the representative methods according to their motivations, structures, and application scenarios. Because of their differences on training settings and dedicated vision tasks, we have also evaluated and compared all these existing visual Transformers under different configurations. Furthermore, we have revealed a series of essential but unexploited aspects that may empower such visual Transformers to stand out from numerous architectures, e.g., slack high-level semantic embeddings to bridge the gap between the visual Transformers and the sequential ones. Finally, three promising research directions are suggested for future investment. We will continue to update the latest articles and their released source codes at https://github.com/liuyang-ict/awesome-visual-transformers.
Accepted by IEEE Transactions on Neural Networks and Learning Systems (TNNLS)
References in corpus (37)
- Sequence to Sequence Learning with Neural Networks
- A Survey on Visual Transformer
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- Transformers in Vision: A Survey
- PVT v2: Improved Baselines with Pyramid Vision Transformer
- Zero-Shot Text-to-Image Generation
- Transformer in Transformer
- BEiT: BERT Pre-Training of Image Transformers
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Conditional Positional Encodings for Vision Transformers
- DeepViT: Towards Deeper Vision Transformer
- LocalViT: Analyzing Locality in Vision Transformers
- Focal Self-attention for Local-Global Interactions in Vision Transformers
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale Up
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- Per-Pixel Classification is Not All You Need for Semantic Segmentation
- Efficient DETR: Improving End-to-End Object Detector with Dense Prior
- ResT: An Efficient Transformer for Visual Recognition
- End-to-End Object Detection with Adaptive Clustering Transformer
- Self-Supervised Learning with Swin Transformers
- Attention-Based Transformers for Instance Segmentation of Cells in Microstructures
- How Much Position Information Do Convolutional Neural Networks Encode?
- Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
- SOLQ: Segmenting Objects by Learning Queries
- ISTR: End-to-End Instance Segmentation with Transformers
- Position, Padding and Predictions: A Deeper Look at Position Information in CNNs
- Vision Transformers with Patch Diversification
- Refiner: Refining Self-attention for Vision Transformers
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet
- MST: Masked Self-Supervised Transformer for Visual Representation
- Scaling Vision with Sparse Mixture of Experts
- VOLO: Vision Outlooker for Visual Recognition
- Transferring Inductive Biases through Knowledge Distillation
- Incorporating Convolution Designs into Visual Transformers