1 paper · 1 filter
Alyssa Unell, Miguel Fuentes, Brenna Li +4
Large language models (LLMs) are increasingly integrated into clinical systems, making it essential to evaluate the real-world utility of these systems. However, static benchmarks…