TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
arXiv:2102.04306
Abstract
Medical image segmentation is an essential prerequisite for developing healthcare systems, especially for disease diagnosis and treatment planning. On various medical image segmentation tasks, the u-shaped architecture, also known as U-Net, has become the de-facto standard and achieved tremendous success. However, due to the intrinsic locality of convolution operations, U-Net generally demonstrates limitations in explicitly modeling long-range dependency. Transformers, designed for sequence-to-sequence prediction, have emerged as alternative architectures with innate global self-attention mechanisms, but can result in limited localization abilities due to insufficient low-level details. In this paper, we propose TransUNet, which merits both Transformers and U-Net, as a strong alternative for medical image segmentation. On one hand, the Transformer encodes tokenized image patches from a convolution neural network (CNN) feature map as the input sequence for extracting global contexts. On the other hand, the decoder upsamples the encoded features which are then combined with the high-resolution CNN feature maps to enable precise localization. We argue that Transformers can serve as strong encoders for medical image segmentation tasks, with the combination of U-Net to enhance finer details by recovering localized spatial information. TransUNet achieves superior performances to various competing methods on different medical applications including multi-organ segmentation and cardiac segmentation. Code and models are available at https://github.com/Beckschen/TransUNet.
13 pages, 3 figures
References in corpus (3)
Cited by in corpus (43)
- A General Survey on Attention Mechanisms in Deep Learning
- SwinNet: Swin Transformer drives edge-aware RGB-D and RGB-T salient object detection
- SCTransNet: Spatial-channel Cross Transformer Network for Infrared Small Target Detection
- CSWin-UNet: Transformer UNet with Cross-Shaped Windows for Medical Image Segmentation
- FCN-Transformer Feature Fusion for Polyp Segmentation
- ViT-V-Net: Vision Transformer for Unsupervised Volumetric Medical Image Registration
- CaraNet: Context Axial Reverse Attention Network for Segmentation of Small Medical Objects
- Boundary-aware Transformers for Skin Lesion Segmentation
- Is it Time to Replace CNNs with Transformers for Medical Images?
- CoTr: Efficiently Bridging CNN and Transformer for 3D Medical Image Segmentation
- Multiscale Vision Transformers
- LeViT-UNet: Make Faster Encoders with Transformer for Medical Image Segmentation
- A Transformer-based Generative Adversarial Network for Brain Tumor Segmentation
- FDiff-Fusion:Denoising diffusion fusion network based on fuzzy learning for 3D medical image segmentation
- HistoSeg : Quick attention with multi-loss function for multi-structure segmentation in digital histology images
- SpecTr: Spectral Transformer for Hyperspectral Pathology Image Segmentation
- Glance-and-Gaze Vision Transformer
- Prototype Learning Guided Hybrid Network for Breast Tumor Segmentation in DCE-MRI
- Transformer-Unet: Raw Image Processing with Unet
- Complex Network for Complex Problems: A comparative study of CNN and Complex-valued CNN
- Medical Image Segmentation Using Squeeze-and-Expansion Transformers
- SKA Science Data Challenge 2: analysis and results
- DnSwin: Toward Real-World Denoising via Continuous Wavelet Sliding-Transformer
- Transformer-Based Source-Free Domain Adaptation
- Federated Split Vision Transformer for COVID-19 CXR Diagnosis using Task-Agnostic Training
- GLIMS: Attention-Guided Lightweight Multi-Scale Hybrid Network for Volumetric Semantic Segmentation
- Automated Measurement of Vascular Calcification in Femoral Endarterectomy Patients Using Deep Learning
- Multi-Compound Transformer for Accurate Biomedical Image Segmentation
- PECI-Net: Bolus segmentation from video fluoroscopic swallowing study images using preprocessing ensemble and cascaded inference
- SDNet: mutil-branch for single image deraining using swin
- A Multi-Branch Hybrid Transformer Networkfor Corneal Endothelial Cell Segmentation
- Advanced Deep Learning Architectures for Accurate Detection of Subsurface Tile Drainage Pipes from Remote Sensing Images
- Transformation Invariant Cancerous Tissue Classification Using Spatially Transformed DenseNet
- Cross-Domain Transfer Learning with CoRTe: Consistent and Reliable Transfer from Black-Box to Lightweight Segmentation Model
- RadioNet: Transformer based Radio Map Prediction Model For Dense Urban Environments
- Combining CNNs With Transformer for Multimodal 3D MRI Brain Tumor Segmentation With Self-Supervised Pretraining
- Dispensed Transformer Network for Unsupervised Domain Adaptation
- Sequential Learning on Liver Tumor Boundary Semantics and Prognostic Biomarker Mining
- ST-DETR: Spatio-Temporal Object Traces Attention Detection Transformer
- UltraPose: Synthesizing Dense Pose with 1 Billion Points by Human-body Decoupling 3D Model
- Class-Incremental Domain Adaptation with Smoothing and Calibration for Surgical Report Generation
- Improving FHB Screening in Wheat Breeding Using an Efficient Transformer Model
- Muscle volume quantification: guiding transformers with anatomical priors