1 paper
Simon Ging, Philipp Arnold, Sebastian Walter +6
Recent 3D CT vision-language models align volumes with reports via contrastive pretraining, but typically rely on limited public data and provide only coarse global supervision. We…