activity
20242026
collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2025

Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks

Krithik Vishwanath, Mrigayu Ghosh, Anton Alyakin +3

Specialized clinical AI assistants are rapidly entering medical practice, often framed as safer or more reliable than general-purpose large language models (LLMs). Yet, unlike fron…

cs.CL2025

Generalist Foundation Models Are Not Clinical Enough for Hospital Operations

Lavender Y. Jiang, Angelica Chen, Xu Han +16

Hospitals and healthcare systems rely on operational decisions that determine patient flow, cost, and quality of care. Despite strong performance on medical knowledge and conversat…

cs.CL2025

MedMobile: A mobile-sized language model with clinical capabilities

Krithik Vishwanath, Jaden Stryker, Anton Alyakin +2

Language models (LMs) have demonstrated expert-level reasoning and recall abilities in medicine. However, computational costs and privacy concerns are mounting barriers to wide-sca…

cs.CL2025

Evaluating the performance and fragility of large language models on the self-assessment for neurological surgeons

Krithik Vishwanath, Anton Alyakin, Mrigayu Ghosh +5

The Congress of Neurological Surgeons Self-Assessment for Neurological Surgeons (CNS-SANS) questions are widely used by neurosurgical residents to prepare for written board examina…

cs.CL2025

Medical large language models are easily distracted

Krithik Vishwanath, Anton Alyakin, Daniel Alexander Alber +3

Large language models (LLMs) have the potential to transform medicine, but real-world clinical scenarios contain extraneous information that can hinder performance. The rise of ass…

cs.CL2025

It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education

Shrutika Singh, Anton Alyakin, Daniel Alexander Alber +9

The performance of Large Language Models (LLMs) on multiple-choice question (MCQ) benchmarks is frequently cited as proof of their medical capabilities. We hypothesized that LLM pe…