Publications (32)
First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations
David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj +54
The paper introduces NOHARM, a benchmark of 1,100 primary‑care to specialist consultation cases, to evaluate how often large language models and retrieval‑augmented clinical AI too…
for Interaction Prediction
David Wu, Yunnan Wu
The 2021 Waymo Interaction Prediction Challenge introduced a problem of predicting the future trajectories and confidences of two interacting agents jointly. We developed a solutio…
Scalable Training of Mixture-of-Experts Models with Megatron Core
Zijie Yan, Hongxiao Bai, Xin Yao +42
Scaling Mixture-of-Experts (MoE) training introduces systems challenges absent in dense models. Because each token activates only a subset of experts, this sparsity allows total pa…
Improving Chess Commentaries by Combining Language Models with Symbolic Reasoning Engines
Andrew Lee, David Wu, Emily Dinan +1
Despite many recent advancements in language modeling, state-of-the-art language models lack grounding in the real world and struggle with tasks involving complex reasoning. Meanwh…
What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models
Payal Chandak, Victoria Alkin, David Wu +11
Medicine is inherently pluralistic. Principles such as autonomy, beneficence, nonmaleficence, and justice routinely conflict, and such ethical dilemmas often sharply divide reasona…
Asking the Right Questions: Benchmarking Large Language Models in the Development of Clinical Consultation Templates
Liam G. McCoy, Fateme Nateghi Haredasht, Kanav Chopra +15
This study evaluates the capacity of large language models (LLMs) to generate structured clinical consultation templates for electronic consultation. Using 145 expert-crafted templ…