5 papers
SL5 Standard for AI Security
Lisa Thiergart, Yoav Tzfati, Peter Wagstaff +3
Security Level 5 (SL5) is a security posture for AI systems that could plausibly thwart top-priority operations by the world's most cyber-capable institutions: those with extensive…
Mechanisms to Verify International Agreements About AI Development
Aaron Scher, Lisa Thiergart
International agreements about AI development may be required to reduce catastrophic risks from advanced AI systems. However, agreements about such a high-stakes technology must be…
What AI evaluations for preventing catastrophic risks can and cannot do
Peter Barnett, Lisa Thiergart
AI evaluations are an important component of the AI governance toolkit, underlying current approaches to safety cases for preventing catastrophic risks. Our paper examines what the…
Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation
Peter Barnett, Lisa Thiergart
As AI systems advance, AI evaluations are becoming an important pillar of regulations for ensuring safety. We argue that such regulation should require developers to explicitly ide…
Steering Language Models With Activation Engineering
Alexander Matt Turner, Lisa Thiergart, Gavin Leech +4
Prompt engineering and finetuning aim to maximize language model performance on a given metric (like toxicity reduction). However, these methods do not fully elicit a model's capab…