A Survey on Visual Transformer
arXiv:2012.12556 · doi:10.1109/TPAMI.2022.3152247
Abstract
Transformer, first applied to the field of natural language processing, is a type of deep neural network mainly based on the self-attention mechanism. Thanks to its strong representation capabilities, researchers are looking at ways to apply transformer to computer vision tasks. In a variety of visual benchmarks, transformer-based models perform similar to or better than other types of networks such as convolutional and recurrent neural networks. Given its high performance and less need for vision-specific inductive bias, transformer is receiving more and more attention from the computer vision community. In this paper, we review these vision transformer models by categorizing them in different tasks and analyzing their advantages and disadvantages. The main categories we explore include the backbone network, high/mid-level vision, low-level vision, and video processing. We also include efficient transformer methods for pushing transformer into real device-based applications. Furthermore, we also take a brief look at the self-attention mechanism in computer vision, as it is the base component in transformer. Toward the end of this paper, we discuss the challenges and provide several further research directions for vision transformers.
Accepted by TPAMI 2022
References in corpus (79)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Learning Transferable Visual Models From Natural Language Supervision
- Language Models are Few-Shot Learners
- PCT: Point cloud transformer
- MLP-Mixer: An all-MLP Architecture for Vision
- Zero-Shot Text-to-Image Generation
- Transformer in Transformer
- Recurrent Models of Visual Attention
- BEiT: BERT Pre-Training of Image Transformers
- Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Conditional Positional Encodings for Vision Transformers
- CogView: Mastering Text-to-Image Generation via Transformers
- TransTrack: Multiple Object Tracking with Transformer
- DeepViT: Towards Deeper Vision Transformer
- A Comprehensive Survey of Scene Graphs: Generation and Application
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Point Transformer
- LocalViT: Analyzing Locality in Vision Transformers
- Reducing Transformer Depth on Demand with Structured Dropout
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale Up
- Diagnose like a Radiologist: Attention Guided Convolutional Neural Network for Thorax Disease Classification
- XCiT: Cross-Covariance Image Transformers
- The Evolved Transformer
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- Understanding and Improving Layer Normalization
- CvT: Introducing Convolutions to Vision Transformers
- Efficient DETR: Improving End-to-End Object Detector with Dense Prior
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Toward Transformer-Based Object Detection
- Lite Transformer with Long-Short Range Attention
- ConvBERT: Improving BERT with Span-based Dynamic Convolution
- Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer
- End-to-End Object Detection with Adaptive Clustering Transformer
- One-Shot Object Detection with Co-Attention and Co-Excitation
- Self-Supervised Learning with Swin Transformers
- Uformer: A General U-Shaped Transformer for Image Restoration
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- Attention-Based Transformers for Instance Segmentation of Cells in Microstructures
- RegionViT: Regional-to-Local Attention for Vision Transformers
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
- Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks
- Efficient Self-supervised Vision Transformers for Representation Learning
- Do You Even Need Attention? A Stack of Feed-Forward Layers Does Surprisingly Well on ImageNet
- On the Computational Efficiency of Training Neural Networks
- TrTr: Visual Tracking with Transformer
- NAT: Neural Architecture Transformer for Accurate and Compact Architectures
- ConvTransformer: A Convolutional Transformer Network for Video Frame Synthesis
- CMT: Convolutional Neural Networks Meet Vision Transformers
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
- Efficient 8-Bit Quantization of Transformer Neural Machine Language Translation Model
- TFPose: Direct Human Pose Estimation with Transformers
- Multiscale Vision Transformers
- SOLQ: Segmenting Objects by Learning Queries
- ISTR: End-to-End Instance Segmentation with Transformers
- Is Attention Better Than Matrix Decomposition?
- Model Rubik's Cube: Twisting Resolution, Depth and Width for TinyNets
- Refiner: Refining Self-attention for Vision Transformers
- DETR for Crowd Pedestrian Detection
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet
- Fully Transformer Networks for Semantic Image Segmentation
- Oriented Object Detection with Transformer
- Spatiotemporal Transformer for Video-based Person Re-identification
- MST: Masked Self-Supervised Transformer for Visual Representation
- A Video Is Worth Three Views: Trigeminal Transformers for Video-based Person Re-identification
- VOLO: Vision Outlooker for Visual Recognition
- Associating Objects with Transformers for Video Object Segmentation
- Visual Parser: Representing Part-whole Hierarchies with Transformers
- SceneFormer: Indoor Scene Generation with Transformers
- Test-Time Personalization with a Transformer for Human Pose Estimation
- Attention is all you need for Videos: Self-attention based Video Summarization using Universal Transformers
- Patch Slimming for Efficient Vision Transformers
- MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens
- Classification by Attention: Scene Graph Classification with Prior Knowledge
- KVT: k-NN Attention for Boosting Vision Transformers
Cited by in corpus (32)
- Transformer in Transformer
- Human Action Recognition from Various Data Modalities: A Review
- Dawn of the transformer era in speech emotion recognition: closing the valence gap
- VLP: A Survey on Vision-Language Pre-training
- Focal Self-attention for Local-Global Interactions in Vision Transformers
- Building extraction with vision transformer
- A Transformer-Based Feature Segmentation and Region Alignment Method For UAV-View Geo-Localization
- Video Transformers: A Survey
- Query2Label: A Simple Transformer Way to Multi-Label Classification
- CDTrans: Cross-domain Transformer for Unsupervised Domain Adaptation
- TFPose: Direct Human Pose Estimation with Transformers
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
- Face Transformer for Recognition
- A Survey of Visual Transformers
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- Collaborative Training of Medical Artificial Intelligence Models with non-uniform Labels
- Soft Sensing Transformer: Hundreds of Sensors are Worth a Single Word
- Predicting Mechanical Properties from Microstructure Images in Fiber-reinforced Polymers using Convolutional Neural Networks
- Video Crowd Localization with Multi-focus Gaussian Neighborhood Attention and a Large-Scale Benchmark
- Rebalanced Zero-shot Learning
- Exploring Advances in Transformers and CNN for Skin Lesion Diagnosis on Small Datasets
- PhysFormer: Facial Video-based Physiological Measurement with Temporal Difference Transformer
- Robust Human Motion Forecasting using Transformer-based Model
- Biologically Inspired Oscillating Activation Functions Can Bridge the Performance Gap between Biological and Artificial Neurons
- Human Image Generation: A Comprehensive Survey
- Hire-MLP: Vision MLP via Hierarchical Rearrangement
- A State-of-the-art Survey of Artificial Neural Networks for Whole-slide Image Analysis:from Popular Convolutional Neural Networks to Potential Visual Transformers
- IMAGO: A family photo album dataset for a socio-historical analysis of the twentieth century
- Transformation Invariant Cancerous Tissue Classification Using Spatially Transformed DenseNet
- Monocular Road Planar Parallax Estimation
- MimicBot: Combining Imitation and Reinforcement Learning to win in Bot Bowl
- Interflow: Aggregating Multi-layer Feature Mappings with Attention Mechanism