PVT v2: Improved Baselines with Pyramid Vision Transformer
arXiv:2106.13797 · doi:10.1007/s41095-022-0274-8
Abstract
Transformer recently has presented encouraging progress in computer vision. In this work, we present new baselines by improving the original Pyramid Vision Transformer (PVT v1) by adding three designs, including (1) linear complexity attention layer, (2) overlapping patch embedding, and (3) convolutional feed-forward network. With these modifications, PVT v2 reduces the computational complexity of PVT v1 to linear and achieves significant improvements on fundamental vision tasks such as classification, detection, and segmentation. Notably, the proposed PVT v2 achieves comparable or better performances than recent works such as Swin Transformer. We hope this work will facilitate state-of-the-art Transformer researches in computer vision. Code is available at https://github.com/whai362/PVT.
Accepted to CVMJ 2022
References in corpus (10)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Transformer in Transformer
- MMDetection: Open MMLab Detection Toolbox and Benchmark
- Twins: Revisiting the Design of Spatial Attention in Vision Transformers
- Conditional Positional Encodings for Vision Transformers
- LocalViT: Analyzing Locality in Vision Transformers
- CvT: Introducing Convolutions to Vision Transformers
- Sparse R-CNN: End-to-End Object Detection with Learnable Proposals
- How Much Position Information Do Convolutional Neural Networks Encode?
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
Cited by in corpus (76)
- Transformers in Vision: A Survey
- MedViT: A Robust Vision Transformer for Generalized Medical Image Classification
- A survey of the Vision Transformers and their CNN-Transformer based Variants
- Polyp-PVT: Polyp Segmentation with Pyramid Vision Transformers
- P2T: Pyramid Pooling Transformer for Scene Understanding
- FCN-Transformer Feature Fusion for Polyp Segmentation
- CBNet: A Composite Backbone Network Architecture for Object Detection
- Video Polyp Segmentation: A Deep Learning Perspective
- MISSFormer: An Effective Medical Image Segmentation Transformer
- Advances in Deep Concealed Scene Understanding
- Masked-attention Mask Transformer for Universal Image Segmentation
- A survey on efficient vision transformers: algorithms, techniques, and performance benchmarking
- Salient Object Detection in Optical Remote Sensing Images Driven by Transformer
- Point Cloud Classification Using Content-based Transformer via Clustering in Feature Space
- RFAConv: Receptive-Field Attention Convolution for Improving Convolutional Neural Networks
- ZoomNeXt: A Unified Collaborative Pyramid Network for Camouflaged Object Detection
- CycleMLP: A MLP-like Architecture for Dense Prediction
- SAM Struggles in Concealed Scenes -- Empirical Study on Segment Anything
- DCN-T: Dual Context Network with Transformer for Hyperspectral Image Classification
- TransXNet: Learning Both Global and Local Dynamics with a Dual Dynamic Token Mixer for Visual Recognition
- A Survey on Deep Learning for Polyp Segmentation: Techniques, Challenges and Future Trends
- LRT: An Efficient Low-Light Restoration Transformer for Dark Light Field Images
- Long-Short Transformer: Efficient Transformers for Language and Vision
- RGBX: Image decomposition and synthesis using material- and lighting-aware diffusion models
- GCoNet+: A Stronger Group Collaborative Co-Salient Object Detector
- A Survey of Visual Transformers
- CrowdFormer: Weakly-supervised Crowd counting with Improved Generalizability
- HAFormer: Unleashing the Power of Hierarchy-Aware Features for Lightweight Semantic Segmentation
- Pixel-Level Clustering Network for Unsupervised Image Segmentation
- S-MLPv2: Improved Spatial-Shift MLP Architecture for Vision
- SwinCross: Cross-modal Swin Transformer for Head-and-Neck Tumor Segmentation in PET/CT Images
- Hierarchical Vision Transformers for Cardiac Ejection Fraction Estimation
- SimCol3D -- 3D Reconstruction during Colonoscopy Challenge
- SpecDETR: A transformer-based hyperspectral point object detection network
- Resolution Enhancement Processing on Low Quality Images Using Swin Transformer Based on Interval Dense Connection Strategy
- Spectrum-driven Mixed-frequency Network for Hyperspectral Salient Object Detection
- Low-Resolution Self-Attention for Semantic Segmentation
- EViT: An Eagle Vision Transformer with Bi-Fovea Self-Attention
- HRPVT: High-Resolution Pyramid Vision Transformer for medium and small-scale human pose estimation
- Vision Transformers: From Semantic Segmentation to Dense Prediction
- Transformers-based architectures for stroke segmentation: A review
- Rethinking Cross-Domain Pedestrian Detection: A Background-Focused Distribution Alignment Framework for Instance-Free One-Stage Detectors
- Dynamic Background Reconstruction via MAE for Infrared Small Target Detection
- UniNeXt: Exploring A Unified Architecture for Vision Recognition
- SeisT: A foundational deep learning model for earthquake monitoring tasks
- SparX: A Sparse Cross-Layer Connection Mechanism for Hierarchical Vision Mamba and Transformer Networks
- A 28nm 0.22μJ/token memory-compute-intensity-aware CNN-Transformer accelerator with hybrid-attention-based layer-fusion and cascaded pruning for semantic-segmentation
- Softmax-free Linear Transformers
- ELMformer: Efficient Raw Image Restoration with a Locally Multiplicative Transformer
- Shunted Self-Attention via Multi-Scale Token Aggregation
- Contrastive Conditional Latent Diffusion for Audio-visual Segmentation
- CAFCT-Net: A CNN-Transformer Hybrid Network with Contextual and Attentional Feature Fusion for Liver Tumor Segmentation
- Cross-modal Cognitive Consensus guided Audio-Visual Segmentation
- Co-Salient Object Detection with Semantic-Level Consensus Extraction and Dispersion
- FAST: Faster Arbitrarily-Shaped Text Detector with Minimalist Kernel Representation
- Hire-MLP: Vision MLP via Hierarchical Rearrangement
- EchoSR: Efficient Context Harnessing for Lightweight Image Super-Resolution
- Frontiers in Intelligent Colonoscopy
- An Efficient Dual-Line Decoder Network with Multi-Scale Convolutional Attention for Multi-organ Segmentation
- Dual-stage Hyperspectral Image Classification Model with Spectral Supertoken
- A Unified Pruning Framework for Vision Transformers
- A Spitting Image: Modular Superpixel Tokenization in Vision Transformers
- Weakly Supervised Camouflaged Object Detection Based on the SAM Model and Mask Guidance
- SC3EF: A Joint Self-Correlation and Cross-Correspondence Estimation Framework for Visible and Thermal Image Registration
- Locally Enhanced Self-Attention: Combining Self-Attention and Convolution as Local and Context Terms
- Full-attention based Neural Architecture Search using Context Auto-regression
- Unveiling Deep Shadows: A Survey and Benchmark on Image and Video Shadow Detection, Removal, and Generation in the Deep Learning Era
- SRMF: A Data Augmentation and Multimodal Fusion Approach for Long-Tail UHR Satellite Image Segmentation
- M&M: Tackling False Positives in Mammography with a Multi-view and Multi-instance Learning Sparse Detector
- MGCR-Net:Multimodal Graph-Conditioned Vision-Language Reconstruction Network for Remote Sensing Change Detection
- Transformers in Unsupervised Structure-from-Motion
- Demystify Transformers & Convolutions in Modern Image Deep Networks
- Beyond Grids: Exploring Elastic Input Sampling for Vision Transformers
- Fine-grained spatial-temporal perception for gas leak segmentation
- Ripple Attention for Visual Perception with Sub-quadratic Complexity
- Beyond MACs: Hardware Efficient Architecture Design for Vision Backbones