Attention Mechanisms in Computer Vision: A Survey
arXiv:2111.07624 · doi:10.1007/s41095-022-0271-y
Abstract
Humans can naturally and effectively find salient regions in complex scenes. Motivated by this observation, attention mechanisms were introduced into computer vision with the aim of imitating this aspect of the human visual system. Such an attention mechanism can be regarded as a dynamic weight adjustment process based on features of the input image. Attention mechanisms have achieved great success in many visual tasks, including image classification, object detection, semantic segmentation, video understanding, image generation, 3D vision, multi-modal tasks and self-supervised learning. In this survey, we provide a comprehensive review of various attention mechanisms in computer vision and categorize them according to approach, such as channel attention, spatial attention, temporal attention and branch attention; a related repository https://github.com/MenghaoGuo/Awesome-Vision-Attentions is dedicated to collecting related work. We also suggest future directions for attention mechanism research.
27 pages, 9 figures
References in corpus (36)
- Transformers in Vision: A Survey
- Language Models are Few-Shot Learners
- PCT: Point cloud transformer
- A Structured Self-attentive Sentence Embedding
- MLP-Mixer: An all-MLP Architecture for Vision
- Going Deeper with Convolutions
- Transformer in Transformer
- Recurrent Models of Visual Attention
- DRAW: A Recurrent Neural Network For Image Generation
- BEiT: BERT Pre-Training of Image Transformers
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- Layer Normalization
- An End-to-End Spatio-Temporal Attention Model for Human Action Recognition from Skeleton Data
- DeepViT: Towards Deeper Vision Transformer
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- Diagnose like a Radiologist: Attention Guided Convolutional Neural Network for Thorax Disease Classification
- Stand-Alone Self-Attention in Vision Models
- GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
- Attentional Pooling for Action Recognition
- Masked Autoencoders Are Scalable Vision Learners
- AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
- Query2Label: A Simple Transformer Way to Multi-Label Classification
- An Empirical Study of Spatial Attention Mechanisms in Deep Networks
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- Multi-Context Attention for Human Pose Estimation
- Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks
- VideoLSTM Convolves, Attends and Flows for Action Recognition
- Single Shot Text Detector with Regional Attention
- Is Attention Better Than Matrix Decomposition?
- Decoupled Spatial-Temporal Transformer for Video Inpainting
- SRM : A Style-based Recalibration Module for Convolutional Neural Networks
- FcaNet: Frequency Channel Attention Networks
- VOLO: Vision Outlooker for Visual Recognition
- GTA: Global Temporal Attention for Video Action Understanding
- FuseFormer: Fusing Fine-Grained Information in Transformers for Video Inpainting
- Sampling Equivariant Self-attention Networks for Object Detection in Aerial Images
Cited by in corpus (32)
- BrainGB: A Benchmark for Brain Network Analysis with Graph Neural Networks
- Collaborative Perception in Autonomous Driving: Methods, Datasets and Challenges
- YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object Detection
- Speck: A Smart event-based Vision Sensor with a low latency 327K Neuron Convolutional Neuronal Network Processing Pipeline
- Engineering Semantic Communication: A Survey
- Deep learning for reconstructing protein structures from cryo-EM density maps: recent advances and future directions
- Quantum Neural Network Classifiers: A Tutorial
- Bridging the Gap: Multi-Level Cross-Modality Joint Alignment for Visible-Infrared Person Re-Identification
- DoubleU-NetPlus: A Novel Attention and Context Guided Dual U-Net with Multi-Scale Residual Feature Fusion Network for Semantic Segmentation of Medical Images
- CD-CTFM: A Lightweight CNN-Transformer Network for Remote Sensing Cloud Detection Fusing Multiscale Features
- Visual Attention-based Self-supervised Absolute Depth Estimation using Geometric Priors in Autonomous Driving
- Deep Learning for Pancreas Segmentation: a Systematic Review
- FlightScope: An Experimental Comparative Review of Aircraft Detection Algorithms in Satellite Imagery
- Low-Resolution Self-Attention for Semantic Segmentation
- Constrained tandem neural network assisted inverse design of metasurfaces for microwave absorption
- Transformers-based architectures for stroke segmentation: A review
- FOOL: Addressing the Downlink Bottleneck in Satellite Computing with Neural Feature Compression
- Engineering Artificial Intelligence: Framework, Challenges, and Future Direction
- Skew Class-balanced Re-weighting for Unbiased Scene Graph Generation
- Artificial Intelligence Model for Tumoral Clinical Decision Support Systems
- YOLO-FEDER FusionNet: A Novel Deep Learning Architecture for Drone Detection
- Enhanced Astronomical Source Classification with Integration of Attention Mechanisms and Vision Transformers
- Advances in Artificial Intelligence: A Review for the Creative Industries
- WaveNets: Wavelet Channel Attention Networks
- Unified Domain Adaptive Semantic Segmentation
- An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval
- Spatiotemporal Object Detection for Improved Aerial Vehicle Detection in Traffic Monitoring
- Guiding Attention in End-to-End Driving Models
- Intelligent Anomaly Detection for Lane Rendering Using Transformer with Self-Supervised Pre-Training and Customized Fine-Tuning
- IgCONDA-PET: Weakly-Supervised PET Anomaly Detection using Implicitly-Guided Attention-Conditional Counterfactual Diffusion Modeling -- a Multi-Center, Multi-Cancer, and Multi-Tracer Study
- A Cloud-Based Hybrid Model for Real-Time Detection of BRTA-Approved Licence Plates Using YOLO Tiny and Haar Cascade
- Visual Attention Graph