works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CR2026

GDM AI Control Roadmap

Mary Phuong, Erik Jenner, Laurent Simon +4

The paper presents the GDM AI Control Roadmap, a framework for internal security against potentially misaligned AI agents, including threat modeling, capability‑based mitigation ti…

cs.CY2026

The 2026 Singapore Consensus on Global AI Safety Research Priorities

Stephen Casper, Oskar Galeev, Yoshua Bengio +117

Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential to embracing AI with confidence. The 202…

cs.CR2026

Phantom Transfer: Data Poisoning can Survive Data-Level Defences

Andrew Draganov, Tolga H. Dur, Anandmayi Bhongade +1

We present a data poisoning attack -- Phantom Transfer -- with the property that, even if you know precisely how the poison was placed into an otherwise benign dataset, you cannot…

cs.AI2025

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38

AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known A…

cs.LG2025

Evaluating Frontier Models for Stealth and Situational Awareness

Mary Phuong, Roland S. Zimmermann, Ziyue Wang +6

Recent work has demonstrated the plausibility of frontier AI models scheming -- knowingly and covertly pursuing an objective misaligned with its developer's intentions. Such behavi…

cs.CY2025

From Stability to Inconsistency: A Study of Moral Preferences in LLMs

Monika Jotautaite, Mary Phuong, Chatrik Singh Mangat +1

As large language models (LLMs) increasingly integrate into our daily lives, it becomes crucial to understand their implicit biases and moral tendencies. To address this, we introd…