3 papers
cs.CV2026
Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology
Yusuf Salcan, Simon Ging, Robin Tibor Schirrmeister +4
We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a large-scale bilingual (German/Engli…
cs.CV2026
Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT
Simon Ging, Philipp Arnold, Sebastian Walter +6
Recent 3D CT vision-language models align volumes with reports via contrastive pretraining, but typically rely on limited public data and provide only coarse global supervision. We…
cs.CV2025
Using Knowledge Graphs to harvest datasets for efficient CLIP model training
Simon Ging, Sebastian Walter, Jelena BratuliÄ +3
Training high-quality CLIP models typically requires enormous datasets, which limits the development of domain-specific models -- especially in areas that even the largest CLIP mod…