Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Towards a Deterministic Math Solver for Clinical Language Models
Felipe Ocampo Osorio, Sebastián Andrés Cajas Ordoñez, Maximin Lange +5
Large language models are unreliable at arithmetic, which is a problem for clinical calculators where a single numerical error changes the recommendation. The standard response is…
cs.AI2026
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Sebastián Andrés Cajas Ordóñez, Agastya Munnangi, Aldo Marzullo +13
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a…
cs.AI2026
Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
Quang Bui, Shlok Jaiswal, Samuel Paik-Heintz +14
Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is rou…