5 papers
CNS-Obsidian: A Neurosurgical Vision-Language Model Built From Scientific Publications
Anton Alyakin, Jaden Stryker, Daniel Alexander Alber +29
General-purpose VLMs demonstrate impressive capabilities, but their opaque training on uncurated internet data poses critical limitations for high-stakes decision-making, such as i…
Generalist Foundation Models Are Not Clinical Enough for Hospital Operations
Lavender Y. Jiang, Angelica Chen, Xu Han +16
Hospitals and healthcare systems rely on operational decisions that determine patient flow, cost, and quality of care. Despite strong performance on medical knowledge and conversat…
Evaluating the performance and fragility of large language models on the self-assessment for neurological surgeons
Krithik Vishwanath, Anton Alyakin, Mrigayu Ghosh +5
The Congress of Neurological Surgeons Self-Assessment for Neurological Surgeons (CNS-SANS) questions are widely used by neurosurgical residents to prepare for written board examina…
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
Shrutika Singh, Anton Alyakin, Daniel Alexander Alber +9
The performance of Large Language Models (LLMs) on multiple-choice question (MCQ) benchmarks is frequently cited as proof of their medical capabilities. We hypothesized that LLM pe…
MedG-KRP: Medical Graph Knowledge Representation Probing
Gabriel R. Rosenbaum, Lavender Yao Jiang, Ivaxi Sheth +11
Large language models (LLMs) have recently emerged as powerful tools, finding many medical applications. LLMs' ability to coalesce vast amounts of information from many sources to…