collaborators

5 papers

cs.CY2026

Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts

Alexander K. Saeri, Jess Graham, Michael Noetel +185

Artificial intelligence poses many risks, ranging from familiar present-day harms to unprecedented and potentially catastrophic ones. Effective risk management requires prioritizat…

cs.CY2026

Measuring AI R&D Automation

Alan Chan, Ranay Padarath, Joe Kwon +2

The automation of AI R&D (AIRDA) could have significant implications, but its extent and ultimate effects remain uncertain. We need empirical data to resolve these uncertainties, b…

cs.CL2026

When Large Language Models are More PersuasiveThan Incentivized Humans, and Why

Philipp Schoenegger, Francesco Salvi, Jiacheng Liu +39

Large Language Models (LLMs) have been shown to be highly persuasive, but when and why they outperform humans is still an open question. We compare the persuasiveness of two LLMs (…

cs.AI2026

Internal Deployment Gaps in AI Regulation

Joe Kwon, Stephen Casper

Frontier AI regulations primarily focus on systems deployed to external users, where deployment is more visible and subject to outside scrutiny. However, high-stakes applications c…

cs.LG2024

Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks

Madeline Brumley, Joe Kwon, David Krueger +2

A key objective of interpretability research on large language models (LLMs) is to develop methods for robustly steering models toward desired behaviors. To this end, two distinct…