12 papers
Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement
Patrick Vossler, Jialin Ouyang, F. Richard Guo +5
Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to estimate and rank the causal…
Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries
Jean Feng, Vishal Patel, Patrick Heagerty +5
Physicians now pose millions of clinical questions to AI tools each week, yet these tools are evaluated largely on hypothetical or exam-style questions, not those actually asked in…
Adaptive auditing of AI systems with anytime-valid guarantees
Siyu Zhou, Patrick Vossler, Venkatesh Sivaraman +2
A major bottleneck in characterizing the failure modes of generative AI systems is the cost and time of annotation and evaluation. Consequently, adaptive testing paradigms have gai…
From Fuzzy to Formal: Scaling Hospital Quality Improvement with AI
Patrick Vossler, Jean Feng, Venkat Sivaraman +9
Hospital Quality Improvement (QI) plays a critical role in optimizing healthcare delivery by translating high-level hospital goals into actionable solutions. A critical step of QI…
LLMs Judging LLMs: A Simplex Perspective
Patrick Vossler, Fan Xia, Yifan Mai +2
Given the challenge of automatically evaluating free-form outputs from large language models (LLMs), an increasingly common solution is to use LLMs themselves as the judging mechan…
More Than "Means to an End": Supporting Reasoning with Transparently Designed AI Data Science Processes
Venkatesh Sivaraman, Patrick Vossler, Adam Perer +2
Generative artificial intelligence (AI) tools can now help people perform complex data science tasks regardless of their expertise. While these tools have great potential to help m…