ResViT: Residual vision transformers for multi-modal medical image synthesis
arXiv:2106.16031 · doi:10.1109/TMI.2022.3167808
Abstract
Generative adversarial models with convolutional neural network (CNN) backbones have recently been established as state-of-the-art in numerous medical image synthesis tasks. However, CNNs are designed to perform local processing with compact filters, and this inductive bias compromises learning of contextual features. Here, we propose a novel generative adversarial approach for medical image synthesis, ResViT, that leverages the contextual sensitivity of vision transformers along with the precision of convolution operators and realism of adversarial learning.} ResViT's generator employs a central bottleneck comprising novel aggregated residual transformer (ART) blocks that synergistically combine residual convolutional and transformer modules. Residual connections in ART blocks promote diversity in captured representations, while a channel compression module distills task-relevant information. A weight sharing strategy is introduced among ART blocks to mitigate computational burden. A unified implementation is introduced to avoid the need to rebuild separate synthesis models for varying source-target modality configurations. Comprehensive demonstrations are performed for synthesizing missing sequences in multi-contrast MRI, and CT images from MRI. Our results indicate superiority of ResViT against competing CNN- and transformer-based methods in terms of qualitative observations and quantitative metrics.
References in corpus (9)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Conditional Generative Adversarial Nets
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- VTGAN: Semi-supervised Retinal Image Synthesis and Disease Prediction using Vision Transformers
- CoTr: Efficiently Bridging CNN and Transformer for 3D Medical Image Segmentation
- PTNet: A High-Resolution Infant MRI Synthesizer Based on Transformer
- GANBERT: Generative Adversarial Networks with Bidirectional Encoder Representations from Transformers for MRI to PET synthesis
- Semi-Supervised Learning of Mutually Accelerated MRI Synthesis without Fully-Sampled Ground Truths
Cited by in corpus (30)
- Adaptive Diffusion Priors for Accelerated MRI Reconstruction
- Transformers in Healthcare: A Survey
- Recent Progress in Transformer-based Medical Image Analysis
- TranSMS: Transformers for Super-Resolution Calibration in Magnetic Particle Imaging
- One Model to Synthesize Them All: Multi-contrast Multi-scale Transformer for Missing Data Imputation
- Is attention all you need in medical image analysis? A review
- Adaptive Latent Diffusion Model for 3D Medical Image to Image Translation: Multi-modal Magnetic Resonance Imaging Study
- Data synthesis and adversarial networks: A review and meta-analysis in cancer imaging
- Class-Aware Adversarial Transformers for Medical Image Segmentation
- Unified Multi-Modal Image Synthesis for Missing Modality Imputation
- Multi-scale Transformer Network with Edge-aware Pre-training for Cross-Modality MR Image Synthesis
- LIT-Former: Linking In-plane and Through-plane Transformers for Simultaneous CT Image Denoising and Deblurring
- An Attentive-based Generative Model for Medical Image Synthesis
- PASTA: Pathology-Aware MRI to PET Cross-Modal Translation with Diffusion Models
- Fusion of Satellite Images and Weather Data with Transformer Networks for Downy Mildew Disease Detection
- Unified Brain MR-Ultrasound Synthesis using Multi-Modal Hierarchical Representations
- FD-Net: An Unsupervised Deep Forward-Distortion Model for Susceptibility Artifact Correction in EPI
- 3D Brain and Heart Volume Generative Models: A Survey
- Weakly Supervised Intracranial Hemorrhage Segmentation using Head-Wise Gradient-Infused Self-Attention Maps from a Swin Transformer in Categorical Learning
- A Densely Interconnected Network for Deep Learning Accelerated MRI
- Joint Self-Supervised and Supervised Contrastive Learning for Multimodal MRI Data: Towards Predicting Abnormal Neurodevelopment
- HAGAN: Hybrid Augmented Generative Adversarial Network for Medical Image Synthesis
- CoNeS: Conditional neural fields with shift modulation for multi-sequence MRI translation
- A Simple and Robust Framework for Cross-Modality Medical Image Segmentation applied to Vision Transformers
- Tumor Synthesis conditioned on Radiomics
- Path and Bone-Contour Regularized Unpaired MRI-to-CT Translation
- Unified Cross-Modal Medical Image Synthesis with Hierarchical Mixture of Product-of-Experts
- Flip Distribution Alignment VAE for Multi-Phase MRI Synthesis
- Quasi-multimodal-based pathophysiological feature learning for retinal disease diagnosis
- Translating MRI to PET through Conditional Diffusion Models with Enhanced Pathology Awareness