2 papers
cs.AI2026
Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System
Alyssa Unell, Miguel Fuentes, Brenna Li +4
Large language models (LLMs) are increasingly integrated into clinical systems, making it essential to evaluate the real-world utility of these systems. However, static benchmarks…
cs.CL2026
CARE: A Conformal Safety Layer for Medical Summarization
Suhana Bedi, Bridget Lin, Anson Y. Zhou +5
Large language models (LLMs) are increasingly used for medical summarization, but their outputs can omit medically important information and introduce unsupported claims. Existing…