Recent Progress in Transformer-based Medical Image Analysis
arXiv:2208.06643 · doi:10.1016/j.compbiomed.2023.107268
Abstract
The transformer is primarily used in the field of natural language processing. Recently, it has been adopted and shows promise in the computer vision (CV) field. Medical image analysis (MIA), as a critical branch of CV, also greatly benefits from this state-of-the-art technique. In this review, we first recap the core component of the transformer, the attention mechanism, and the detailed structures of the transformer. After that, we depict the recent progress of the transformer in the field of MIA. We organize the applications in a sequence of different tasks, including classification, segmentation, captioning, registration, detection, enhancement, localization, and synthesis. The mainstream classification and segmentation tasks are further divided into eleven medical image modalities. A large number of experiments studied in this review illustrate that the transformer-based method outperforms existing methods through comparisons with multiple evaluation metrics. Finally, we discuss the open challenges and future opportunities in this field. This task-modality review with the latest contents, detailed information, and comprehensive comparison may greatly benefit the broad MIA community.
Computers in Biology and Medicine Accepted
References in corpus (28)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Rethinking Atrous Convolution for Semantic Image Segmentation
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- Bootstrap your own latent: A new approach to self-supervised Learning
- A Structured Self-attentive Sentence Embedding
- The Medical Segmentation Decathlon
- MedMNIST v2 -- A large-scale lightweight benchmark for 2D and 3D biomedical image classification
- Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)
- Training Generative Adversarial Networks with Limited Data
- Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation
- Linformer: Self-Attention with Linear Complexity
- REFUGE Challenge: A Unified Framework for Evaluating Automated Methods for Glaucoma Assessment from Fundus Photographs
- COVID-19 Image Data Collection
- TransMorph: Transformer for unsupervised medical image registration
- A large annotated medical image dataset for the development and evaluation of segmentation algorithms
- Reliable Tuberculosis Detection using Chest X-ray with Deep Learning, Segmentation and Visualization
- ResViT: Residual vision transformers for multi-modal medical image synthesis
- Med3D: Transfer Learning for 3D Medical Image Analysis
- Escaping the Big Data Paradigm with Compact Transformers
- PanNuke Dataset Extension, Insights and Baselines
- BS-Net: learning COVID-19 pneumonia severity on a large Chest X-Ray dataset
- BIMCV COVID-19+: a large annotated dataset of RX and CT images from COVID-19 patients
- VTGAN: Semi-supervised Retinal Image Synthesis and Disease Prediction using Vision Transformers
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- Transformers in Medical Image Analysis: A Review
- A Transformer-based Generative Adversarial Network for Brain Tumor Segmentation
- Federated Split Vision Transformer for COVID-19 CXR Diagnosis using Task-Agnostic Training
- Data-Efficient Vision Transformers for Multi-Label Disease Classification on Chest Radiographs