UNETR: Transformers for 3D Medical Image Segmentation
arXiv:2103.10504
Abstract
Fully Convolutional Neural Networks (FCNNs) with contracting and expanding paths have shown prominence for the majority of medical image segmentation applications since the past decade. In FCNNs, the encoder plays an integral role by learning both global and local features and contextual representations which can be utilized for semantic output prediction by the decoder. Despite their success, the locality of convolutional layers in FCNNs, limits the capability of learning long-range spatial dependencies. Inspired by the recent success of transformers for Natural Language Processing (NLP) in long-range sequence learning, we reformulate the task of volumetric (3D) medical image segmentation as a sequence-to-sequence prediction problem. We introduce a novel architecture, dubbed as UNEt TRansformers (UNETR), that utilizes a transformer as the encoder to learn sequence representations of the input volume and effectively capture the global multi-scale information, while also following the successful "U-shaped" network design for the encoder and decoder. The transformer encoder is directly connected to a decoder via skip connections at different resolutions to compute the final semantic segmentation output. We have validated the performance of our method on the Multi Atlas Labeling Beyond The Cranial Vault (BTCV) dataset for multi-organ segmentation and the Medical Segmentation Decathlon (MSD) dataset for brain tumor and spleen segmentation tasks. Our benchmarks demonstrate new state-of-the-art performance on the BTCV leaderboard. Code: https://monai.io/research/unetr
Accepted to IEEE Winter Conference on Applications of Computer Vision (WACV) 2022
References in corpus (9)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- BEiT: BERT Pre-Training of Image Transformers
- A large annotated medical image dataset for the development and evaluation of segmentation algorithms
- On the Compactness, Efficiency, and Representation of 3D Convolutional Networks: Brain Parcellation as a Pretext Task
- TransFuse: Fusing Transformers and CNNs for Medical Image Segmentation
- CoTr: Efficiently Bridging CNN and Transformer for 3D Medical Image Segmentation
- LambdaNetworks: Modeling Long-Range Interactions Without Attention
Cited by in corpus (14)
- Recent advances and clinical applications of deep learning in medical image analysis
- Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation
- MISSFormer: An Effective Medical Image Segmentation Transformer
- BiTr-Unet: a CNN-Transformer Combined Network for MRI Brain Tumor Segmentation
- BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once
- GaNDLF: A Generally Nuanced Deep Learning Framework for Scalable End-to-End Clinical Workflows in Medical Imaging
- Axial multi-layer perceptron architecture for automatic segmentation of choroid plexus in multiple sclerosis
- SeqSeg: Learning Local Segments for Automatic Vascular Model Construction
- A Comprehensive Study on Medical Image Segmentation using Deep Neural Networks
- Hepatic vessel segmentation based on 3D swin-transformer with inductive biased multi-head self-attention
- More than Encoder: Introducing Transformer Decoder to Upsample
- Vision Transformer for Classification of Breast Ultrasound Images
- Multi-task learning for joint weakly-supervised segmentation and aortic arch anomaly classification in fetal cardiac MRI
- Calibrated Self-supervised Vision Transformers Improve Intracranial Arterial Calcification Segmentation from Clinical CT Head Scans