collaborators

12 papers

cs.CL2026

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

Tianming Du, Peijie Yu, Sihan Shang +12

The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR sys…

cs.CV2026

EnTrust: Modeling Inter-Modal Conflict for Trustworthy Multimodal Medical Image Analysis

Dwarikanath Mahapatra, Abhijit Das, Behzad Bozorgtabar +5

Multimodal medical imaging fuses complementary anatomical and functional information, yet modalities frequently disagree in pathologically heterogeneous regions. Current segmentati…

cs.CV2026

Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology

Yusuf Salcan, Simon Ging, Robin Tibor Schirrmeister +4

We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a large-scale bilingual (German/Engli…

cs.CV2026

VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs

Govinda Kolli, Adinath Madhavrao Dukre, Behzad Bozorgtabar +2

Medical Vision-Language Models (VLMs) often hallucinate by generating responses based on language priors rather than visual evidence, posing risks in clinical applications. We prop…

cs.CV2026

Refining 3D Medical Segmentation with Verbal Instruction

Kangxian Xie, Jiancheng Yang, Nandor Pinter +3

Accurate 3D anatomical segmentation is essential for clinical diagnosis and surgical planning. However, automated models frequently generate suboptimal shape predictions due to fac…

cs.CV2026

Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT

Simon Ging, Philipp Arnold, Sebastian Walter +6

Recent 3D CT vision-language models align volumes with reports via contrastive pretraining, but typically rely on limited public data and provide only coarse global supervision. We…