4 papers
HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks
Suhana Bedi, Ryan Welch, Ethan Steinberg +12
Healthcare administration accounts for over $1 trillion in annual spending, making it a promising target for LLM-based computer-use agents (CUAs). While clinical applications of LL…
Structured Prompts Improve Evaluation of Language Models
Asad Aali, Muhammad Ahmed Mohsin, Vasiliki Bikia +15
As language models (LMs) are increasingly adopted across domains, high-quality benchmarking frameworks are essential for guiding deployment decisions. In practice, however, framewo…
Monitoring Deployed AI Systems in Health Care
Timothy Keyes, Alison Callahan, Abby S. Pandya +18
Post-deployment monitoring of artificial intelligence (AI) systems in health care is essential to ensure their safety, quality, and sustained benefit-and to support governance deci…
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Suhana Bedi, Hejie Cui, Miguel Fuentes +78
While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinica…