3 papers
cs.LG2026
CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations
Yusuf Kesmen, Aniruddha Mukherjee, Yena Chang +7
Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinica…
cs.CL2026
Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
Fay Elhassan, David Sasu, Alexandra Kulinkina +2
We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive O…
cs.LG2026
MoBayes: A Modular Bayesian Framework for Separating Reasoning from Language in Conversational Clinical Decision Support
Yusuf Kesmen, Fay Elhassan, Jiayi Ma +7
Large language models (LLMs) are increasingly used for conversational clinical decision support, yet they conflate next token prediction with probabilistic decision making. We argu…