9 papers
Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability
Alyssa Unell, Natalie Dullerud, Naomi Boneh +4
LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation. However, the reliability of these judges depends critically on their alignme…
Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System
Alyssa Unell, Miguel Fuentes, Brenna Li +4
Large language models (LLMs) are increasingly integrated into clinical systems, making it essential to evaluate the real-world utility of these systems. However, static benchmarks…
Adoption and Use of LLMs at an Academic Medical Center
Nigam H. Shah, Nerissa Ambers, Abby Pandya +55
While large language models (LLMs) can support clinical documentation needs, standalone tools struggle with "workflow friction" from manual data entry. We developed ChatEHR, a syst…
Structured Prompts Improve Evaluation of Language Models
Asad Aali, Muhammad Ahmed Mohsin, Vasiliki Bikia +15
As language models (LMs) are increasingly adopted across domains, high-quality benchmarking frameworks are essential for guiding deployment decisions. In practice, however, framewo…
Not Another EHR: Reimagining Physician Information Needs with Generative AI Technology
Ruican Zhong, Jiachen Li, Gary Hsieh +15
Electronic health records (EHRs) have improved data accessibility but have also introduced cognitive burden for physicians, given the sheer volume and complexity of the data involv…
DISCO: A Browser-Based Privacy-Preserving Framework for Distributed Collaborative Learning
Julien T. T. Vignoud, Valérian Rousset, Hugo El Guedj +28
Data is often impractical to share for a range of well considered reasons, such as concerns over privacy, intellectual property, and legal constraints. This not only fragments the…