Publications (9)
Open Technical Problems in Open-Weight AI Model Risk Management
Stephen Casper, Kyle O'Brien, Shayne Longpre +19
Frontier AI models with openly available weights are steadily becoming more powerful and widely adopted. However, compared to proprietary models, open-weight models pose different…
Auditing Work: Exploring the New York City algorithmic bias audit regime
Lara Groves, Jacob Metcalf, Alayna Kennedy +2
In July 2023, New York City (NYC) initiated the first algorithm auditing system for commercial machine-learning systems. Local Law 144 (LL 144) mandates NYC-based employers using a…
The Reality of AI and Biorisk
Aidan Peppin, Anka Reuel, Stephen Casper +10
To accurately and confidently answer the question 'could an AI model or system increase biorisk', it is necessary to have both a sound theoretical threat model for how AI models or…
Measuring and mitigating overreliance to build human-compatible AI
Lujain Ibrahim, Katherine M. Collins, Sunnie S. Y. Kim +14
Large language models (LLMs) distinguish themselves from previous technologies by functioning as collaborative ``thought partners,'' capable of engaging more fluidly in natural lan…
Going public: the role of public participation approaches in commercial AI labs
Lara Groves, Aidan Peppin, Andrew Strait +1
In recent years, discussions of responsible AI practices have seen growing support for "participatory AI" approaches, intended to involve members of the public in the design and de…
A Multi-Turn Framework for Evaluating AI Misuse in Fraud and Cybercrime Scenarios
Kimberly T. Mai, Anna Gausen, Magda Dubois +5
AI is increasingly being used to assist fraud and cybercrime. However, it is unclear the extent to which current large language models can provide useful information for complex cr…