8 papers
Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology
Yusuf Salcan, Simon Ging, Robin Tibor Schirrmeister +4
We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a large-scale bilingual (German/Engli…
Eliciting associations between clinical variables from LLMs via comparison questions across populations
Fabian Kabus, Kian Kordtomeikel, Thomas Brox +3
The training data of large language models (LLMs) comprises a wide range of biomedical literature, reflecting data from many different patient populations. We investigate how it mi…
Assessing Multimodal Chronic Wound Embeddings with Expert Triplet Agreement
Fabian Kabus, Julia Hindel, Jelena BratuliÄ +6
Recessive dystrophic epidermolysis bullosa (RDEB) is a rare genetic skin disorder for which clinicians greatly benefit from finding similar cases using images and clinical text. Ho…
On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
Jelena BratuliÄ, Sudhanshu Mittal, Thomas Brox +1
Feed-forward 3D reconstruction models such as DUSt3R, VGGT, and Depth Anything 3 (DA3) are transformer-based foundation models that infer camera geometry and dense scene structure…
Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT
Simon Ging, Philipp Arnold, Sebastian Walter +6
Recent 3D CT vision-language models align volumes with reports via contrastive pretraining, but typically rely on limited public data and provide only coarse global supervision. We…
Unlocking In-Context Learning for Natural Datasets Beyond Language Modelling
Jelena BratuliÄ, Sudhanshu Mittal, David T. Hoffmann +5
Large Language Models (LLMs) exhibit In-Context Learning (ICL), which enables the model to perform new tasks conditioning only on the examples provided in the context without updat…