11 citations · 29 across the 3 of their papers we have counts for
3 papers
cs.CY2024★ 9 cited
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
Seliem El-Sayed, Canfer Akbulut, Amanda McCroskery +17
Recent generative AI systems have demonstrated more advanced persuasive capabilities and are increasingly permeating areas of life where they can influence decision-making. Generat…
cs.LG2024★ 11 cited
Evaluating Frontier Models for Dangerous Capabilities
Mary Phuong, Matthew Aitchison, Elliot Catt +24
To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evalu…
cs.AI2022★ 9 cited
Structured access: an emerging paradigm for safe AI deployment
Toby Shevlane
Structured access is an emerging paradigm for the safe deployment of artificial intelligence (AI). Instead of openly disseminating AI systems, developers facilitate controlled, arm…