2 papers
cs.LG2026
Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models
Bruce A. Bassett, Amy Rouillard, Sitwala Mundia +8
Background: Large language models (LLMs) are increasingly proposed for diagnostic support, but few evaluations use real-world multimodal inpatient data, particularly in low and mid…
cs.LG2026
Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
Amy Rouillard, Sitwala Mundia, Linda Camara +8
Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators. Here, we evaluate an…