Transformers in Vision: A Survey
arXiv:2101.01169 · doi:10.1145/3505244
Abstract
Astounding results from Transformer models on natural language tasks have intrigued the vision community to study their application to computer vision problems. Among their salient benefits, Transformers enable modeling long dependencies between input sequence elements and support parallel processing of sequence as compared to recurrent networks e.g., Long short-term memory (LSTM). Different from convolutional networks, Transformers require minimal inductive biases for their design and are naturally suited as set-functions. Furthermore, the straightforward design of Transformers allows processing multiple modalities (e.g., images, videos, text and speech) using similar processing blocks and demonstrates excellent scalability to very large capacity networks and huge datasets. These strengths have led to exciting progress on a number of vision tasks using Transformer networks. This survey aims to provide a comprehensive overview of the Transformer models in the computer vision discipline. We start with an introduction to fundamental concepts behind the success of Transformers i.e., self-attention, large-scale pre-training, and bidirectional encoding. We then cover extensive applications of transformers in vision including popular recognition tasks (e.g., image classification, object detection, action recognition, and segmentation), generative modeling, multi-modal tasks (e.g., visual-question answering, visual reasoning, and visual grounding), video processing (e.g., activity recognition, video forecasting), low-level vision (e.g., image super-resolution, image enhancement, and colorization) and 3D analysis (e.g., point cloud classification and segmentation). We compare the respective advantages and limitations of popular techniques both in terms of architectural design and their experimental value. Finally, we provide an analysis on open research directions and possible future works.
30 pages (Accepted in ACM Computing Surveys December 2021)
References in corpus (68)
- Distilling the Knowledge in a Neural Network
- Semi-Supervised Classification with Graph Convolutional Networks
- Learning Transferable Visual Models From Natural Language Supervision
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Bootstrap your own latent: A new approach to self-supervised Learning
- Language Models are Few-Shot Learners
- The Kinetics Human Action Video Dataset
- PVT v2: Improved Baselines with Pyramid Vision Transformer
- NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding
- Is Space-Time Attention All You Need for Video Understanding?
- Transformer in Transformer
- Linformer: Self-Attention with Linear Complexity
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- Twins: Revisiting the Design of Spatial Attention in Vision Transformers
- Generating Long Sequences with Sparse Transformers
- Conditional Positional Encodings for Vision Transformers
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Axial Attention in Multidimensional Transformers
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- DeepViT: Towards Deeper Vision Transformer
- Reformer: The Efficient Transformer
- Intriguing Properties of Vision Transformers
- Escaping the Big Data Paradigm with Compact Transformers
- LocalViT: Analyzing Locality in Vision Transformers
- Focal Self-attention for Local-Global Interactions in Vision Transformers
- Spatial Temporal Transformer Network for Skeleton-based Action Recognition
- XCiT: Cross-Covariance Image Transformers
- Perceiver IO: A General Architecture for Structured Inputs & Outputs
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- Restormer: Efficient Transformer for High-Resolution Image Restoration
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- CvT: Introducing Convolutions to Vision Transformers
- ResT: An Efficient Transformer for Visual Recognition
- All Tokens Matter: Token Labeling for Training Better Vision Transformers
- Pix2seq: A Language Modeling Framework for Object Detection
- Perceiver: General Perception with Iterative Attention
- Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer
- Pre-Trained Image Processing Transformer
- Random Feature Attention
- Self-Supervised Learning with Swin Transformers
- Uformer: A General U-Shaped Transformer for Image Restoration
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- Attention-Based Transformers for Instance Segmentation of Cells in Microstructures
- TransReID: Transformer-based Object Re-Identification
- RegionViT: Regional-to-Local Attention for Vision Transformers
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
- Efficient Self-supervised Vision Transformers for Representation Learning
- Do You Even Need Attention? A Stack of Feed-Forward Layers Does Surprisingly Well on ImageNet
- Sparse Sinkhorn Attention
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning
- Referring Transformer: A One-step Approach to Multi-task Visual Grounding
- Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
- On the Turing Completeness of Modern Neural Network Architectures
- Multiscale Vision Transformers
- A Universal Representation Transformer Layer for Few-Shot Image Classification
- LambdaNetworks: Modeling Long-Range Interactions Without Attention
- End-to-End Human Pose and Mesh Reconstruction with Transformers
- Parameter Efficient Multimodal Transformers for Video Representation Learning
- AutoFormer: Searching Transformers for Visual Recognition
- SceneFormer: Indoor Scene Generation with Transformers
- Deep Amortized Clustering
- Topological Planning with Transformers for Vision-and-Language Navigation
- On Improving Adversarial Transferability of Vision Transformers
- Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding
- Attention, please! A survey of Neural Attention Models in Deep Learning
- GLiT: Neural Architecture Search for Global and Local Image Transformer
- Long-Short Temporal Contrastive Learning of Video Transformers
Cited by in corpus (89)
- Attention Mechanisms in Computer Vision: A Survey
- Deep Neural Networks and Tabular Data: A Survey
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- A Survey of Human-in-the-loop for Machine Learning
- Human Action Recognition from Various Data Modalities: A Review
- Deep Learning for Time Series Anomaly Detection: A Survey
- Multimodal Fusion Transformer for Remote Sensing Image Classification
- Intriguing Properties of Vision Transformers
- Combining EfficientNet and Vision Transformers for Video Deepfake Detection
- Focal Self-attention for Local-Global Interactions in Vision Transformers
- Building extraction with vision transformer
- Vision Transformers For Weeds and Crops Classification Of High Resolution UAV Images
- Video Transformers: A Survey
- Short and Long Range Relation Based Spatio-Temporal Transformer for Micro-Expression Recognition
- Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing Images
- A Practical Survey on Faster and Lighter Transformers
- How to avoid machine learning pitfalls: a guide for academic researchers
- Query2Label: A Simple Transformer Way to Multi-Label Classification
- Self-Supervised Masked Convolutional Transformer Block for Anomaly Detection
- Deep Contextual Video Compression
- Backbones-Review: Feature Extraction Networks for Deep Learning and Deep Reinforcement Learning Approaches
- CDTrans: Cross-domain Transformer for Unsupervised Domain Adaptation
- Is it Time to Replace CNNs with Transformers for Medical Images?
- TransReID: Transformer-based Object Re-Identification
- VTGAN: Semi-supervised Retinal Image Synthesis and Disease Prediction using Vision Transformers
- A review on vision-based analysis for automatic dietary assessment
- Self-Contrastive Learning with Hard Negative Sampling for Self-supervised Point Cloud Learning
- Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks
- A review of Generative Adversarial Networks for Electronic Health Records: applications, evaluation measures and data sources
- Swin-transformer-yolov5 For Real-time Wine Grape Bunch Detection
- Multi-Stage Progressive Image Restoration
- ISTR: End-to-End Instance Segmentation with Transformers
- Patch Similarity Aware Data-Free Quantization for Vision Transformers
- ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis
- Neuromorphic Camera Denoising using Graph Neural Network-driven Transformers
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
- A Survey of Visual Transformers
- Creativity and Machine Learning: A Survey
- Dynamic Neural Networks: A Survey
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- Transformer-based Models to Deal with Heterogeneous Environments in Human Activity Recognition
- Are Neural Language Models Good Plagiarists? A Benchmark for Neural Paraphrase Detection
- Characterization of anomalous diffusion through convolutional transformers
- TransformerFusion: Monocular RGB Scene Reconstruction using Transformers
- Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video
- Semantic Labeling of High Resolution Images Using EfficientUNets and Transformers
- RPLHR-CT Dataset and Transformer Baseline for Volumetric Super-Resolution from CT Scans
- MPSN: Motion-aware Pseudo Siamese Network for Indoor Video Head Detection in Buildings
- Multi-Head Self-Attention via Vision Transformer for Zero-Shot Learning
- MalBERT: Using Transformers for Cybersecurity and Malicious Software Detection
- Medical Image Segmentation using LeViT-UNet++: A Case Study on GI Tract Data
- Multi-Scale Hybrid Vision Transformer for Learning Gastric Histology: AI-Based Decision Support System for Gastric Cancer Treatment
- Federated Split Vision Transformer for COVID-19 CXR Diagnosis using Task-Agnostic Training
- Exploring Advances in Transformers and CNN for Skin Lesion Diagnosis on Small Datasets
- Unsupervised Person Re-Identification: A Systematic Survey of Challenges and Solutions
- One-Step Abductive Multi-Target Learning with Diverse Noisy Samples and Its Application to Tumour Segmentation for Breast Cancer
- AutoFormer: Searching Transformers for Visual Recognition
- Rethinking Cooking State Recognition with Vision Transformers
- Detection of Illicit Drug Trafficking Events on Instagram: A Deep Multimodal Multilabel Learning Approach
- IntFormer: Predicting pedestrian intention with the aid of the Transformer architecture
- Morphological Classification of Radio Galaxies with wGAN-supported Augmentation
- PhysFormer: Facial Video-based Physiological Measurement with Temporal Difference Transformer
- PU-Transformer: Point Cloud Upsampling Transformer
- Trans-SVNet: Accurate Phase Recognition from Surgical Videos via Hybrid Embedding Aggregation Transformer
- On Improving Adversarial Transferability of Vision Transformers
- Forensic License Plate Recognition with Compression-Informed Transformers
- End-to-End Trainable Multi-Instance Pose Estimation with Transformers
- Deep neural networks approach to microbial colony detection -- a comparative analysis
- Literature review on vulnerability detection using NLP technology
- Vision Transformer using Low-level Chest X-ray Feature Corpus for COVID-19 Diagnosis and Severity Quantification
- Human Image Generation: A Comprehensive Survey
- Performance Evaluation of Action Recognition Models on Low Quality Videos
- DAF:re: A Challenging, Crowd-Sourced, Large-Scale, Long-Tailed Dataset For Anime Character Recognition
- Prototype Memory for Large-scale Face Representation Learning
- Self-supervised Video-centralised Transformer for Video Face Clustering
- Perceptual Image Quality Assessment with Transformers
- Boosting Few-shot Semantic Segmentation with Transformers
- Semi-Supervised Wide-Angle Portraits Correction by Multi-Scale Transformer
- Investigating Attention Mechanism in 3D Point Cloud Object Detection
- Video Transformer for Deepfake Detection with Incremental Learning
- Visual Framing of Science Conspiracy Videos: Integrating Machine Learning with Communication Theories to Study the Use of Color and Brightness
- A Survey of Fish Tracking Techniques Based on Computer Vision
- PnP-3D: A Plug-and-Play for 3D Point Clouds
- Point Cloud Learning with Transformer
- When Liebig's Barrel Meets Facial Landmark Detection: A Practical Model
- Sequential Random Network for Fine-grained Image Classification
- Advancing biological super-resolution microscopy through deep learning: a brief review
- Transformers for prompt-level EMA non-response prediction
- Survey of Visual-Semantic Embedding Methods for Zero-Shot Image Retrieval