8 papers
Structure-Adaptive Sparse Diffusion in Voxel Space for 3D Medical Image Enhancement
Hongxu Jiang, Fei Li, Boxiao Yu +4
Three-dimensional (3D) medical image enhancement, including denoising and super-resolution, is critical for clinical diagnosis in CT, PET, and MRI. Although diffusion models have s…
Beyond the Embedding Bottleneck: Adaptive Retrieval-Augmented 3D CT Report Generation
Renjie Liang, Yiling Ma, Yang Xing +6
Automated radiology report generation from 3D CT volumes often suffers from incomplete pathology coverage. We provide empirical evidence that this limitation stems from a represent…
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
Yang Xing, Jiong Wu, Savas Ozdemir +4
Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (…
TauGenNet: Plasma-Driven Tau PET Image Synthesis via Text-Guided 3D Diffusion Models
Yuxin Gong, Se-in Jang, Wei Shao +2
Accurate quantification of tau pathology via tau positron emission tomography (PET) scan is crucial for diagnosing and monitoring Alzheimer's disease (AD). However, the high cost a…
CDPDNet: Integrating Text Guidance with Hybrid Vision Encoders for Medical Image Segmentation
Jiong Wu, Yang Xing, Boxiao Yu +2
Most publicly available medical segmentation datasets are only partially labeled, with annotations provided for a subset of anatomical structures. When multiple datasets are combin…
Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis
Yu Xin, Gorkem Can Ates, Kuang Gong +1
Vision-language models (VLMs) have shown promise in 2D medical image analysis, but extending them to 3D remains challenging due to the high computational demands of volumetric data…