From the 1 of 21 linked papers with an AI index.
21 papers
Symbal: Detecting Systematic Misalignments in Model-Generated Captions
Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier +2
The paper presents Symbal, a dual‑stage method that uses off‑the‑shelf foundation models to automatically detect systematic misalignments—recurring caption errors tied to specific…
CheXmix: Unified Generative Pretraining for Vision Language Models in Medical Imaging
Ashwin Kumar, Robbie Holland, Corey Barrett +8
Recent medical multimodal foundation models are built as multimodal LLMs (MLLMs) by connecting a CLIP-pretrained vision encoder to an LLM using LLaVA-style finetuning. This two-sta…
A Reasoning-Enabled Vision-Language Foundation Model for Chest X-ray Interpretation
Yabin Zhang, Chong Wang, Yunhe Gao +19
Chest X-rays (CXRs) are among the most frequently performed imaging examinations worldwide, yet rising imaging volumes increase radiologist workload and the risk of diagnostic erro…
Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Models
Yabin Zhang, Maya Varma, Yunhe Gao +4
Out-of-distribution (OOD) detection aims to identify samples that deviate from in-distribution (ID). One popular pipeline addresses this by introducing negative labels distant from…
Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision
Yunhe Gao, Yabin Zhang, Chong Wang +5
Foundation models have transformed vision and language by learning general-purpose representations from large-scale unlabeled data, yet 3D medical imaging lacks analogous approache…
Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset
Louis Blankemeier, Ashwin Kumar, Joseph Paul Cohen +37
The large volume of abdominal computed tomography (CT) scans coupled with the shortage of radiologists have intensified the need for automated medical image analysis tools. Previou…