5 papers
Measuring and mitigating overreliance to build human-compatible AI
Lujain Ibrahim, Katherine M. Collins, Sunnie S. Y. Kim +14
Large language models (LLMs) distinguish themselves from previous technologies by functioning as collaborative ``thought partners,'' capable of engaging more fluidly in natural lan…
A Multi-Turn Framework for Evaluating AI Misuse in Fraud and Cybercrime Scenarios
Kimberly T. Mai, Anna Gausen, Magda Dubois +5
AI is increasingly being used to assist fraud and cybercrime. However, it is unclear the extent to which current large language models can provide useful information for complex cr…
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
Christopher Summerfield, Lennart Luettgau, Magda Dubois +9
We examine recent research that asks whether current AI systems may be developing a capacity for "scheming" (covertly and strategically pursuing misaligned goals). We compare curre…
Towards Best Practices for Open Datasets for LLM Training
Stefan Baack, Stella Biderman, Kasia Odrozek +36
Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in…
The Reality of AI and Biorisk
Aidan Peppin, Anka Reuel, Stephen Casper +10
To accurately and confidently answer the question 'could an AI model or system increase biorisk', it is necessary to have both a sound theoretical threat model for how AI models or…