papers

Publications (56)

cs.CV2023

Measuring the Success of Diffusion Models at Imitating Human Artists

Stephen Casper, Zifan Guo, Shreya Mogulothu +5

Modern diffusion models have set the state-of-the-art in AI image generation. Their success is due, in part, to training on Internet-scale data which often includes copyrighted wor…

cs.AI2026

Internal Deployment Gaps in AI Regulation

Joe Kwon, Stephen Casper

Frontier AI regulations primarily focus on systems deployed to external users, where deployment is more visible and subject to outside scrutiny. However, high-stakes applications c…

cs.CY2025

International AI Safety Report 2025: Second Key Update: Technical Safeguards and Risk Management

Yoshua Bengio, Stephen Clare, Carina Prunkl +66

This second update to the 2025 International AI Safety Report assesses new developments in general-purpose AI risk management over the past year. It examines how researchers, publi…

cs.CY2025

Open Problems in Technical AI Governance

Anka Reuel, Ben Bucknall, Stephen Casper +30

AI progress is creating a growing range of risks and opportunities, but it is often unclear how they should be navigated. In many cases, the barriers and uncertainties faced are at…

cs.AI2026

The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence

Peter Slattery, Alexander K. Saeri, Emily A. C. Grundy +7

Artificial intelligence (AI) is reshaping society, from video generation to medical diagnosis, coding agents to autonomous vehicles. Yet researchers, policymakers, and technology c…

cs.LG2023

Diagnostics for Deep Neural Networks with Automated Copy/Paste Attacks

Stephen Casper, Kaivalya Hariharan, Dylan Hadfield-Menell

This paper considers the problem of helping humans exercise scalable oversight over deep neural networks (DNNs). Adversarial examples can be useful by helping to reveal weaknesses…

cs.AI2023

Red Teaming with Mind Reading: White-Box Adversarial Policies Against RL Agents

Stephen Casper, Taylor Killian, Gabriel Kreiman +1

Adversarial examples can be useful for identifying vulnerabilities in AI systems before they are deployed. In reinforcement learning (RL), adversarial policies can be developed by…

cs.CY2026

Open Weight AI Models Require Proportional Evaluation Approaches

Patricia Paskov, Christopher Rodriguez, Sunishchal Dev +1

Open-weight AI models (OWMs), or models released with publicly-available weights, are distributing rapidly and approaching the performance levels of leading closed-weight AI models…

cs.CL2024

Eight Methods to Evaluate Robust Unlearning in LLMs

Aengus Lynch, Phillip Guo, Aidan Ewart +2

Machine unlearning can be useful for removing harmful capabilities and memorized text from large language models (LLMs), but there are not yet standardized methods for rigorously e…

cs.LG2022

Quantifying Local Specialization in Deep Neural Networks

Shlomi Hod, Daniel Filan, Stephen Casper +2

A neural network is locally specialized to the extent that parts of its computational graph (i.e. structure) can be abstractly represented as performing some comprehensible sub-tas…

cs.CL2023

Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Rusheb Shah, Quentin Feuillade--Montixi, Soroush Pour +3

Despite efforts to align large language models to produce harmless responses, they are still vulnerable to jailbreak prompts that elicit unrestricted behaviour. In this work, we in…

cs.CY2026

Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies

Miles Brundage, Noemi Dreksler, Aidan Homewood +45

We outline a vision for frontier AI auditing, which we define as rigorous third-party verification of frontier AI developers' safety and security claims, and evaluation of their sy…

cs.LG2025

Open Problems in Machine Unlearning for AI Safety

Fazl Barez, Tingchen Fu, Ameya Prabhu +16

As AI systems become more capable, widely deployed, and increasingly autonomous in critical areas such as cybersecurity, biological research, and healthcare, ensuring their safety…

cs.AI2025

The Singapore Consensus on Global AI Safety Research Priorities

Yoshua Bengio, Tegan Maharaj, Luke Ong +84

Rapidly improving AI capabilities and autonomy hold significant promise of transformation, but are also driving vigorous debate on how to ensure that AI is safe, i.e., trustworthy,…

cs.CY2025

Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs

Ariba Khan, Stephen Casper, Dylan Hadfield-Menell

Research on the 'cultural alignment' of Large Language Models (LLMs) has emerged in response to growing interest in understanding representation across diverse stakeholders. Curren…

cs.CY2026

International AI Safety Report 2026

Yoshua Bengio, Stephen Clare, Carina Prunkl +89

The International AI Safety Report 2026 synthesises the current scientific evidence on the capabilities, emerging risks, and safety of general-purpose AI systems. The report series…

cs.CY2025

Pitfalls of Evidence-Based AI Policy

Stephen Casper, David Krueger, Dylan Hadfield-Menell

Nations across the world are working to govern AI. However, from a technical perspective, there is uncertainty and disagreement on the best way to do this. Meanwhile, recent debate…

cs.CR2025

Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Stephen Casper, Lennart Schulze, Oam Patel +1

Despite extensive diagnostics and debugging by developers, AI systems sometimes exhibit harmful unintended behaviors. Finding and fixing these is challenging because the attack sur…

cs.CY2026

Legal Alignment for Safe and Ethical AI

Noam Kolt, Nicholas Caputo, Jack Boeglin +14

Alignment of artificial intelligence (AI) encompasses the normative problem of specifying how AI systems should act and the technical problem of ensuring AI systems comply with tho…

cs.CL2026

STACK: Adversarial Attacks on LLM Safeguard Pipelines

Ian R. McKenzie, Oskar J. Hollinsworth, Tom Tseng +5

Frontier AI developers are relying on layers of safeguards to protect against catastrophic misuse of AI systems. Anthropic and OpenAI guard their latest Opus 4 model and GPT-5 mode…

cs.LG2025

Open Problems in Mechanistic Interpretability

Lee Sharkey, Bilal Chughtai, Joshua Batson +26

Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goa…

cs.AI2023

Achilles Heels for AGI/ASI via Decision Theoretic Adversaries

Stephen Casper

As progress in AI continues to advance, it is important to know how advanced systems will make choices and in what ways they may fail. Machines can already outsmart humans in some…

cs.LG2024

Rethinking Machine Unlearning for Large Language Models

Sijia Liu, Yuanshun Yao, Jinghan Jia +11

We explore machine unlearning (MU) in the domain of large language models (LLMs), referred to as LLM unlearning. This initiative aims to eliminate undesirable data influence (e.g.,…

cs.LG2023

Red Teaming Deep Neural Networks with Feature Synthesis Tools

Stephen Casper, Yuxiao Li, Jiawei Li +4

Interpretable AI tools are often motivated by the goal of understanding model behavior in out-of-distribution (OOD) contexts. Despite the attention this area of study receives, the…

cs.CY2026

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems

Leon Staufer, Kevin Feng, Kevin Wei +6

Agentic AI systems are increasingly capable of performing professional and personal tasks with limited human involvement. However, tracking these developments is difficult because…

cs.CY2024

Black-Box Access is Insufficient for Rigorous AI Audits

Stephen Casper, Carson Ezell, Charlotte Siegmann +18

External audits of AI systems are increasingly recognized as a key mechanism for AI governance. The effectiveness of an audit, however, depends on the degree of access granted to a…

cs.CL2020

Probing Neural Dialog Models for Conversational Understanding

Abdelrhman Saleh, Tovly Deutsch, Stephen Casper +2

The predominant approach to open-domain dialog generation relies on end-to-end training of neural models on chat datasets. However, this approach provides little insight as to what…

cs.CY2026

Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts

Alexander K. Saeri, Jess Graham, Michael Noetel +185

Artificial intelligence poses many risks, ranging from familiar present-day harms to unprecedented and potentially catastrophic ones. Effective risk management requires prioritizat…

cs.CR2025

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Zora Che, Stephen Casper, Robert Kirk +12

Evaluations of large language model (LLM) risks and capabilities are increasingly being incorporated into AI risk management and governance frameworks. Currently, most risk evaluat…

cs.NE2021

Clusterability in Neural Networks

Daniel Filan, Stephen Casper, Shlomi Hod +3

The learned weights of a neural network have often been considered devoid of scrutable internal structure. In this paper, however, we look for structure in the form of clusterabili…

cs.CY2026

Open Technical Problems in Open-Weight AI Model Risk Management

Stephen Casper, Kyle O'Brien, Shayne Longpre +19

Frontier AI models with openly available weights are steadily becoming more powerful and widely adopted. However, compared to proprietary models, open-weight models pose different…

cs.CL2023

Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?

Kevin Liu, Stephen Casper, Dylan Hadfield-Menell +1

Neural language models (LMs) can be used to evaluate the truth of factual statements in two ways: they can be either queried for statement probabilities, or probed for internal rep…

cs.CL2023

Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Stephen Casper, Jason Lin, Joe Kwon +2

Deploying large language models (LMs) can pose hazards from harmful outputs such as toxic or false text. Prior work has introduced automated tools that elicit harmful outputs to id…

cs.LG2026

Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs

Kyle O'Brien, Stephen Casper, Quentin Anthony +7

Open-weight AI systems offer unique benefits, including enhanced transparency, open research, and decentralized access. However, they are vulnerable to tampering attacks which can…

cs.CY2025

International AI Safety Report

Yoshua Bengio, Sören Mindermann, Daniel Privitera +93

The first International AI Safety Report comprehensively synthesizes the current evidence on the capabilities, risks, and safety of advanced AI systems. The report was mandated by…

cs.LG2021

Frivolous Units: Wider Networks Are Not Really That Wide

Stephen Casper, Xavier Boix, Vanessa D'Amario +4

A remarkable characteristic of overparameterized deep neural networks (DNNs) is that their accuracy does not degrade when the network's width is increased. Recent evidence suggests…

cs.LG2024

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Usman Anwar, Abulhair Saparov, Javier Rando +39

This work identifies 18 foundational challenges in assuring the alignment and safety of large language models (LLMs). These challenges are organized into three different categories…

cs.CY2026

Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns

Max Kamachee, Stephen Casper, Michelle L. Ding +4

In 2022, AI image generators crossed a threshold, enabling much more efficient and dynamic production of photorealistic deepfake images than before. This enabled opportunities for…

cs.AI2025

Practical Principles for AI Cost and Compute Accounting

Stephen Casper, Luke Bailey, Tim Schreier

Policymakers increasingly use development cost and compute as proxies for AI capabilities and risks. Recent laws have introduced regulatory requirements for models or developers th…

cs.LG2023

Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks

Tilman Räuker, Anson Ho, Stephen Casper +1

The last decade of machine learning has seen drastic increases in scale and capabilities. Deep neural networks (DNNs) are increasingly being deployed in the real world. However, th…

cs.CY2026

Expanding External Access To Frontier AI Models For Dangerous Capability Evaluations

Jacob Charnock, Alejandro Tlaie, Kyle O'Brien +2

Frontier AI companies increasingly rely on external evaluations to assess risks from dangerous capabilities before deployment. However, external evaluators often receive limited mo…

cs.CY2026

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack

Cristian Trout, Sanmi Koyejo, Sasha Romanosky +34

The paper proposes a comprehensive AI insurance framework to price and manage risks from the emerging AI agent economy, outlining an eight‑component stack for data collection, mode…

#ai insurance#risk assessment#catastrophe modeling#frontier ai
cs.CR2026

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

Saad Hossain, Tom Tseng, Punya Syon Pandey +8

As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, be…

cs.AI2023

Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Stephen Casper, Xander Davies, Claudia Shi +29

Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals. RLHF has emerged as the central method used to finetune state-of…

cs.CY2025

Audit Cards: Contextualizing AI Evaluations

Leon Staufer, Mick Yang, Anka Reuel +1

AI governance frameworks increasingly rely on audits, yet the results of their underlying evaluations require interpretation and context to be meaningfully informative. Even techni…

cs.CY2025

International Scientific Report on the Safety of Advanced AI (Interim Report)

Yoshua Bengio, Sören Mindermann, Daniel Privitera +41

This is the interim publication of the first International Scientific Report on the Safety of Advanced AI. The report synthesises the scientific understanding of general-purpose AI…

cs.LG2026

Open Problems in Frontier AI Risk Management

Marta Ziosi, Miro Plueckebaum, Stephen Casper +26

Frontier AI both amplifies existing risks and introduces qualitatively novel challenges. Not only is there a notable lack of stable scientific consensus resulting from the rapid pa…

cs.LG2025

Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Abhay Sheshadri, Aidan Ewart, Phillip Guo +8

Large language models (LLMs) can often be made to behave in undesirable ways that they are explicitly fine-tuned not to. For example, the LLM red-teaming literature has produced a…

cs.AI2025

The Reality of AI and Biorisk

Aidan Peppin, Anka Reuel, Stephen Casper +10

To accurately and confidently answer the question 'could an AI model or system increase biorisk', it is necessary to have both a sound theoretical threat model for how AI models or…

cs.AI2024

Multilevel Interpretability Of Artificial Neural Networks: Leveraging Framework And Methods From Neuroscience

Zhonghao He, Jascha Achterberg, Katie Collins +13

As deep learning systems are scaled up to many billions of parameters, relating their internal structure to external behaviors becomes very challenging. Although daunting, this pro…

cs.LG2024

The SaTML '24 CNN Interpretability Competition: New Innovations for Concept-Level Interpretability

Stephen Casper, Jieun Yun, Joonhyuk Baek +13

Interpretability techniques are valuable for helping humans understand and oversee AI systems. The SaTML 2024 CNN Interpretability Competition solicited novel methods for studying…

cs.LG2025

Obfuscated Activations Bypass LLM Latent-Space Defenses

Luke Bailey, Alex Serrano, Abhay Sheshadri +7

Recent latent-space monitoring techniques have shown promise as defenses against LLM attacks. These defenses act as scanners that seek to detect harmful activations before they lea…

cs.CR2025

What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks

Nathalie Kirch, Constantin Weisser, Severin Field +2

Jailbreaks have been a central focus of research regarding the safety and reliability of large language models (LLMs), yet the mechanisms underlying these attacks remain poorly und…

cs.SE2025

The AI Agent Index

Stephen Casper, Luke Bailey, Rosco Hunter +12

Leading AI developers and startups are increasingly deploying agentic AI systems that can plan and execute complex tasks with limited human involvement. However, there is currently…

cs.LG2025

Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives

Leo Schwinn, Yan Scholten, Tom Wollschläger +4

Misaligned research objectives have considerably hindered progress in adversarial robustness research over the past decade. For instance, an extensive focus on optimizing target me…

cs.LG2023

Robust Feature-Level Adversaries are Interpretability Tools

Stephen Casper, Max Nadeau, Dylan Hadfield-Menell +1

The literature on adversarial attacks in computer vision typically focuses on pixel-level perturbations. These tend to be very difficult to interpret. Recent work that manipulates…