From the 1 of 16 linked papers with an AI index.
16 papers
Symbal: Detecting Systematic Misalignments in Model-Generated Captions
Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier +2
The paper presents Symbal, a dual‑stage method that uses off‑the‑shelf foundation models to automatically detect systematic misalignments—recurring caption errors tied to specific…
CheXTemporal: A Dataset for Temporally-Grounded Reasoning in Chest Radiography
Eva Prakash, Yunhe Gao, Chong Wang +10
Chest radiograph interpretation requires temporal reasoning over prior and current studies, yet most vision-language models are trained on static image-report pairs and lack explic…
A Reasoning-Enabled Vision-Language Foundation Model for Chest X-ray Interpretation
Yabin Zhang, Chong Wang, Yunhe Gao +19
Chest X-rays (CXRs) are among the most frequently performed imaging examinations worldwide, yet rising imaging volumes increase radiologist workload and the risk of diagnostic erro…
Clinician input steers AI toward accurate and harmful recommendations
Ivan Lopez, Selin S. Everett, Bryan J. Bunning +10
Large language models (LLMs) are entering clinical workflows, yet evaluations rarely assess how clinician reasoning shapes model behavior during clinical interactions. Using 61 cur…
Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset
Louis Blankemeier, Ashwin Kumar, Joseph Paul Cohen +37
The large volume of abdominal computed tomography (CT) scans coupled with the shortage of radiologists have intensified the need for automated medical image analysis tools. Previou…
Attention Head Entropy of LLMs Predicts Answer Correctness
Sophie Ostmeier, Brian Axelrod, Maya Varma +6
Large language models (LLMs) often generate plausible yet incorrect answers, posing risks in safety-critical settings such as medicine. Human evaluation is expensive, and LLM-as-ju…