Publications (13)
Adoption and Use of LLMs at an Academic Medical Center
Nigam H. Shah, Nerissa Ambers, Abby Pandya +55
While large language models (LLMs) can support clinical documentation needs, standalone tools struggle with "workflow friction" from manual data entry. We developed ChatEHR, a syst…
Simple Agents Outperform Experts in Biomedical Imaging Workflow Optimization
Xuefei, Wang, Kai A. Horstmann +9
Adapting production-level computer vision tools to bespoke scientific datasets is a critical "last mile" bottleneck. Current solutions are impractical: fine-tuning requires large a…
Federated Learning Enables Big Data for Rare Cancer Boundary Detection
Sarthak Pati, Ujjwal Baid, Brandon Edwards +276
Although machine learning (ML) has shown promise in numerous domains, there are concerns about generalizability to out-of-sample data. This is currently addressed by centrally shar…
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Suhana Bedi, Hejie Cui, Miguel Fuentes +78
While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinica…
Reasoning and Tools for Human-Level Forecasting
Elvis Hsieh, Preston Fu, Jonathan Chen
Language models (LMs) trained on web-scale datasets are largely successful due to their ability to memorize large amounts of training data, even if only present in a few examples.…
Multi-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource Languages
Chukwuebuka Anyaegbuna, Eduardo Juan Perez Guerrero, Jerry Liu +6
Language barriers affect 27.3 million U.S. residents with non-English language preference, yet professional medical translation remains costly and often unavailable. We evaluated f…
Constrained Design of a Binary Instrument in a Partially Linear Model
Tim Morrison, Minh Nguyen, Jonathan Chen +2
We study the question of how best to assign an encouragement in a randomized encouragement study. In our setting, units arrive with covariates, receive a nudge toward treatment or…
Standing on FURM ground -- A framework for evaluating Fair, Useful, and Reliable AI Models in healthcare systems
Alison Callahan, Duncan McElfresh, Juan M. Banda +21
The impact of using artificial intelligence (AI) to guide patient care or operational processes is an interplay of the AI model's output, the decision-making protocol based on that…
Explainable PCGML via Game Design Patterns
Matthew Guzdial, Joshua Reno, Jonathan Chen +2
Procedural content generation via Machine Learning (PCGML) is the umbrella term for approaches that generate content for games via machine learning. One of the benefits of PCGML is…
Superhuman performance of a large language model on the reasoning tasks of a physician
Peter G. Brodeur, Thomas A. Buckley, Zahir Kanjee +22
A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing sy…
Bias-preserving and error-detectable entangling operations in a superconducting dual-rail system
Nitish Mehta, James D. Teoh, Taewan Noh +69
For useful quantum computation, error-corrected machines are required that can dramatically reduce the inevitable errors experienced by physical qubits. While significant progress…
Demonstrating a superconducting dual-rail cavity qubit with erasure-detected logical measurements
Kevin S. Chou, Tali Shemma, Heather McCarrick +31
A critical challenge in developing scalable error-corrected quantum systems is the accumulation of errors while performing operations and measurements. One promising approach is to…
Friend, Collaborator, Student, Manager: How Design of an AI-Driven Game Level Editor Affects Creators
Matthew Guzdial, Nicholas Liao, Jonathan Chen +6
Machine learning advances have afforded an increase in algorithms capable of creating art, music, stories, games, and more. However, it is not yet well-understood how machine learn…