From the 1 of 6 linked papers with an AI index.
6 papers
GDM AI Control Roadmap
Mary Phuong, Erik Jenner, Laurent Simon +4
The paper presents the GDM AI Control Roadmap, a framework for internal security against potentially misaligned AI agents, including threat modeling, capability‑based mitigation ti…
The 2026 Singapore Consensus on Global AI Safety Research Priorities
Stephen Casper, Oskar Galeev, Yoshua Bengio +117
Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential to embracing AI with confidence. The 202…
Phantom Transfer: Data Poisoning can Survive Data-Level Defences
Andrew Draganov, Tolga H. Dur, Anandmayi Bhongade +1
We present a data poisoning attack -- Phantom Transfer -- with the property that, even if you know precisely how the poison was placed into an otherwise benign dataset, you cannot…
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38
AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known A…
Evaluating Frontier Models for Stealth and Situational Awareness
Mary Phuong, Roland S. Zimmermann, Ziyue Wang +6
Recent work has demonstrated the plausibility of frontier AI models scheming -- knowingly and covertly pursuing an objective misaligned with its developer's intentions. Such behavi…
From Stability to Inconsistency: A Study of Moral Preferences in LLMs
Monika Jotautaite, Mary Phuong, Chatrik Singh Mangat +1
As large language models (LLMs) increasingly integrate into our daily lives, it becomes crucial to understand their implicit biases and moral tendencies. To address this, we introd…