4 papers
Towards a Deterministic Math Solver for Clinical Language Models
Felipe Ocampo Osorio, Sebastián Andrés Cajas Ordoñez, Maximin Lange +5
Large language models are unreliable at arithmetic, which is a problem for clinical calculators where a single numerical error changes the recommendation. The standard response is…
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Sebastián Andrés Cajas Ordóñez, Agastya Munnangi, Aldo Marzullo +13
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a…
Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
Quang Bui, Shlok Jaiswal, Samuel Paik-Heintz +14
Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is rou…
Analyzing Diversity in Healthcare LLM Research: A Scientometric Perspective
David Restrepo, Chenwei Wu, Constanza Vásquez-Venegas +5
The deployment of large language models (LLMs) in healthcare has demonstrated substantial potential for enhancing clinical decision-making, administrative efficiency, and patient o…