TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model
arXiv:2607.13812 · doi:10.1609/aaai.v39i21.34433
Abstract
We introduce TCAM-Diff, a novel 3D medical image generation model that reduces the memory requirements to encode and generate high-resolution 3D data. This model utilizes a decoder-only autoencoder method to learn triplane representation from dense volume and leverages generalization operations to prevent overfitting. Subsequently, it uses a triplane-aware cross-attention diffusion model to learn and integrate these features effectively. Furthermore, the features generated by the diffusion model can be rapidly transformed into 3D volumes using a pre-trained decoder module. Our experiments on three different scales of medical datasets, BrainTumour 128 x 128 x 128, Pancreas 256 x 256 x 256, and Colon 512 x 512 x 512, demonstrate outstanding results. We utilized MSE and SSIM to assess reconstruction quality and leveraged the Wasserstein Generative Adversarial Network (W-GAN) critic to assess generative quality. Comparisons with existing approaches show that our method gives better reconstruction and generation results than other encoder-decoder methods with similar-sized latent spaces.
Accepted at AAAI 2025. Code is available at https://github.com/Fredy-Zhang/TCAM-Diff
References in corpus (6)
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- High-Resolution Image Synthesis with Latent Diffusion Models
- A large annotated medical image dataset for the development and evaluation of segmentation algorithms
- Brain Imaging Generation with Latent Diffusion Models
- 3D-StyleGAN: A Style-Based Generative Adversarial Network for Generative Modeling of Three-Dimensional Medical Images
- 3D Neural Field Generation using Triplane Diffusion