Publications (19)
How well do LLMs cite relevant medical references? An evaluation framework and analyses
Kevin Wu, Eric Wu, Ally Cassasola +7
Large language models (LLMs) are currently being used to answer medical questions across a variety of clinical domains. Recent top-performing commercial LLMs, in particular, are al…
Infrastructure for AI Agents
Alan Chan, Kevin Wei, Sihao Huang +5
AI agents plan and execute interactions in open-ended environments. For example, OpenAI's Operator can use a web browser to do product comparisons and buy online goods. Much resear…
RCTs for Frontier AI Governance: Methodological Challenges and Solutions for Human Uplift Studies
Patricia Paskov, Kevin Wei, Shen Zhou Hong +7
Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or similar methodologies, increasingly inform fro…
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
Cristian Trout, Sanmi Koyejo, Sasha Romanosky +34
The paper proposes a comprehensive AI insurance framework to price and manage risks from the emerging AI agent economy, outlining an eight‑component stack for data collection, mode…
Local US officials' views on the impacts and governance of AI: Evidence from 2022 and 2023 survey waves
Sophia Hatz, Noemi Dreksler, Kevin Wei +1
This paper presents a survey of local US policymakers' views on the future impact and regulation of AI. Our survey provides insight into US policymakers' expectations regarding the…
Measurement of charged-pion production in deep-inelastic scattering off nuclei with the CLAS detector
S. Moran, R. Dupre, H. Hakobyan +141
Background: Energetic quarks in nuclear DIS propagate through the nuclear medium. Processes that are believed to occur inside nuclei include quark energy loss through medium-stimul…