Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025★ 1 cited
Generalist Foundation Models Are Not Clinical Enough for Hospital Operations
Lavender Y. Jiang, Angelica Chen, Xu Han +16
Hospitals and healthcare systems rely on operational decisions that determine patient flow, cost, and quality of care. Despite strong performance on medical knowledge and conversat…
cs.CL2025
Evaluating the performance and fragility of large language models on the self-assessment for neurological surgeons
Krithik Vishwanath, Anton Alyakin, Mrigayu Ghosh +5
The Congress of Neurological Surgeons Self-Assessment for Neurological Surgeons (CNS-SANS) questions are widely used by neurosurgical residents to prepare for written board examina…
cs.CL2025
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
Shrutika Singh, Anton Alyakin, Daniel Alexander Alber +9
The performance of Large Language Models (LLMs) on multiple-choice question (MCQ) benchmarks is frequently cited as proof of their medical capabilities. We hypothesized that LLM pe…