medical imaging

TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model

arXiv:2607.13812 · doi:10.1609/aaai.v39i21.34433

summary

The paper presents TCAM-Diff, a diffusion-based model that uses a triplane-aware cross‑attention mechanism and a decoder‑only autoencoder to efficiently generate high‑resolution 3D medical images with reduced memory usage.

Abstract

We introduce TCAM-Diff, a novel 3D medical image generation model that reduces the memory requirements to encode and generate high-resolution 3D data. This model utilizes a decoder-only autoencoder method to learn triplane representation from dense volume and leverages generalization operations to prevent overfitting. Subsequently, it uses a triplane-aware cross-attention diffusion model to learn and integrate these features effectively. Furthermore, the features generated by the diffusion model can be rapidly transformed into 3D volumes using a pre-trained decoder module. Our experiments on three different scales of medical datasets, BrainTumour 128 x 128 x 128, Pancreas 256 x 256 x 256, and Colon 512 x 512 x 512, demonstrate outstanding results. We utilized MSE and SSIM to assess reconstruction quality and leveraged the Wasserstein Generative Adversarial Network (W-GAN) critic to assess generative quality. Comparisons with existing approaches show that our method gives better reconstruction and generation results than other encoder-decoder methods with similar-sized latent spaces.

Accepted at AAAI 2025. Code is available at https://github.com/Fredy-Zhang/TCAM-Diff

Topics & keywords

#3d image generation#diffusion models#triplane representation#cross-attention#medical imagingtriplane-aware cross-attentiondecoder-only autoencoderWasserstein GAN criticMSESSIM
TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model · wovepaper