One Model to Synthesize Them All: Multi-contrast Multi-scale Transformer for Missing Data Imputation
arXiv:2204.13738 · doi:10.1109/TMI.2023.3261707
Abstract
Multi-contrast magnetic resonance imaging (MRI) is widely used in clinical practice as each contrast provides complementary information. However, the availability of each imaging contrast may vary amongst patients, which poses challenges to radiologists and automated image analysis algorithms. A general approach for tackling this problem is missing data imputation, which aims to synthesize the missing contrasts from existing ones. While several convolutional neural networks (CNN) based algorithms have been proposed, they suffer from the fundamental limitations of CNN models, such as the requirement for fixed numbers of input and output channels, the inability to capture long-range dependencies, and the lack of interpretability. In this work, we formulate missing data imputation as a sequence-to-sequence learning problem and propose a multi-contrast multi-scale Transformer (MMT), which can take any subset of input contrasts and synthesize those that are missing. MMT consists of a multi-scale Transformer encoder that builds hierarchical representations of inputs combined with a multi-scale Transformer decoder that generates the outputs in a coarse-to-fine fashion. The proposed multi-contrast Swin Transformer blocks can efficiently capture intra- and inter-contrast dependencies for accurate image synthesis. Moreover, MMT is inherently interpretable as it allows us to understand the importance of each input contrast in different regions by analyzing the in-built attention maps of Transformer blocks in the decoder. Extensive experiments on two large-scale multi-contrast MRI datasets demonstrate that MMT outperforms the state-of-the-art methods quantitatively and qualitatively.
IEEE TMI accepted final version
References in corpus (7)
- Understanding Neural Networks Through Deep Visualization
- Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation
- ResViT: Residual vision transformers for multi-modal medical image synthesis
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale Up
- ViT-V-Net: Vision Transformer for Unsupervised Volumetric Medical Image Registration
- VTGAN: Semi-supervised Retinal Image Synthesis and Disease Prediction using Vision Transformers
- PTNet: A High-Resolution Infant MRI Synthesizer Based on Transformer
Cited by in corpus (5)
- Transformers in Healthcare: A Survey
- Similarity and Quality Metrics for MR Image-To-Image Translation
- Unified Multi-Modal Image Synthesis for Missing Modality Imputation
- Multi-scale Transformer Network with Edge-aware Pre-training for Cross-Modality MR Image Synthesis
- CoNeS: Conditional neural fields with shift modulation for multi-sequence MRI translation