Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation
arXiv:2105.05537
Abstract
In the past few years, convolutional neural networks (CNNs) have achieved milestones in medical image analysis. Especially, the deep neural networks based on U-shaped architecture and skip-connections have been widely applied in a variety of medical image tasks. However, although CNN has achieved excellent performance, it cannot learn global and long-range semantic information interaction well due to the locality of the convolution operation. In this paper, we propose Swin-Unet, which is an Unet-like pure Transformer for medical image segmentation. The tokenized image patches are fed into the Transformer-based U-shaped Encoder-Decoder architecture with skip-connections for local-global semantic feature learning. Specifically, we use hierarchical Swin Transformer with shifted windows as the encoder to extract context features. And a symmetric Swin Transformer-based decoder with patch expanding layer is designed to perform the up-sampling operation to restore the spatial resolution of the feature maps. Under the direct down-sampling and up-sampling of the inputs and outputs by 4x, experiments on multi-organ and cardiac segmentation tasks demonstrate that the pure Transformer-based U-shaped Encoder-Decoder network outperforms those methods with full-convolution or the combination of transformer and convolution. The codes and trained models will be publicly available at https://github.com/HuCaoFighting/Swin-Unet.
a drafted manuscript
References in corpus (3)
Cited by in corpus (42)
- Recent advances and clinical applications of deep learning in medical image analysis
- Video-SwinUNet: Spatio-temporal Deep Learning Framework for VFSS Instance Segmentation
- nnFormer: Interleaved Transformer for Volumetric Segmentation
- Modality specific U-Net variants for biomedical image segmentation: A survey
- SUNet: Swin Transformer UNet for Image Denoising
- A Survey on Deep Learning for Skin Lesion Segmentation
- WORD: A large scale dataset, benchmark and clinical applicable study for abdominal organ segmentation from CT image
- Transformers in Healthcare: A Survey
- MISSFormer: An Effective Medical Image Segmentation Transformer
- Recent Progress in Transformer-based Medical Image Analysis
- One Model to Synthesize Them All: Multi-contrast Multi-scale Transformer for Missing Data Imputation
- BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once
- SwinIR: Image Restoration Using Swin Transformer
- Multi-scale Transformer Network with Edge-aware Pre-training for Cross-Modality MR Image Synthesis
- LRT: An Efficient Low-Light Restoration Transformer for Dark Light Field Images
- You Only Train Once: Learning a General Anomaly Enhancement Network with Random Masks for Hyperspectral Anomaly Detection
- Factorizer: A Scalable Interpretable Approach to Context Modeling for Medical Image Segmentation
- LeViT-UNet: Make Faster Encoders with Transformer for Medical Image Segmentation
- Compete to Win: Enhancing Pseudo Labels for Barely-supervised Medical Image Segmentation
- Application of belief functions to medical image segmentation: A review
- TreeFormer: a Semi-Supervised Transformer-based Framework for Tree Counting from a Single High Resolution Image
- Spectral-wise Implicit Neural Representation for Hyperspectral Image Reconstruction
- CiT-Net: Convolutional Neural Networks Hand in Hand with Vision Transformers for Medical Image Segmentation
- Semantic Labeling of High Resolution Images Using EfficientUNets and Transformers
- A Convolutional Vision Transformer for Semantic Segmentation of Side-Scan Sonar Data
- STB-VMM: Swin Transformer Based Video Motion Magnification
- DuDoTrans: Dual-Domain Transformer Provides More Attention for Sinogram Restoration in Sparse-View CT Reconstruction
- Multiclass Segmentation using Teeth Attention Modules for Dental X-ray Images
- Automated Measurement of Vascular Calcification in Femoral Endarterectomy Patients Using Deep Learning
- Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image Reconstruction
- Enabling Large Batch Size Training for DNN Models Beyond the Memory Limit While Maintaining Performance
- Sam2Rad: A Segmentation Model for Medical Images with Learnable Prompts
- Comparative and Interpretative Analysis of CNN and Transformer Models in Predicting Wildfire Spread Using Remote Sensing Data
- Deep Evidential Learning for Radiotherapy Dose Prediction
- Hepatic vessel segmentation based on 3D swin-transformer with inductive biased multi-head self-attention
- An Efficient Dual-Line Decoder Network with Multi-Scale Convolutional Attention for Multi-organ Segmentation
- BiSeg-SAM: Weakly-Supervised Post-Processing Framework for Boosting Binary Segmentation in Segment Anything Models
- Semi-Supervised Wide-Angle Portraits Correction by Multi-Scale Transformer
- Transformation Invariant Cancerous Tissue Classification Using Spatially Transformed DenseNet
- TED-net: Convolution-free T2T Vision Transformer-based Encoder-decoder Dilation network for Low-dose CT Denoising
- Dispensed Transformer Network for Unsupervised Domain Adaptation
- SFB-net for cardiac segmentation: Bridging the semantic gap with attention