Publications (19)
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud +1340
Phantom Transfer: Data Poisoning can Survive Data-Level Defences
Andrew Draganov, Tolga H. Dur, Anandmayi Bhongade +1
Model evaluation for extreme risks
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel +18
Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors
Oliver Makins, Orazio Angelini, Zohreh Shams +1
Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
Rohin Shah, Vikrant Varma, Ramana Kumar +4
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
Benjamin Arnav, Pablo Bernabeu-Pérez, Nathan Helm-Burger +3
Evaluating Frontier Models for Dangerous Capabilities
Mary Phuong, Matthew Aitchison, Elliot Catt +24
Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives
Chloe Li, Mary Phuong, Daniel Tan
Towards Understanding Knowledge Distillation
Mary Phuong, Christoph H. Lampert
From Stability to Inconsistency: A Study of Moral Preferences in LLMs
Monika Jotautaite, Mary Phuong, Chatrik Singh Mangat +1
Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalization
Frank Xiao, Mary Phuong
Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents
Frank Xiao, Mary Phuong
Evaluating Frontier Models for Stealth and Situational Awareness
Mary Phuong, Roland S. Zimmermann, Ziyue Wang +6
Formal Algorithms for Transformers
Mary Phuong, Marcus Hutter
GDM AI Control Roadmap
Mary Phuong, Erik Jenner, Laurent Simon +4
The paper presents the GDM AI Control Roadmap, a framework for internal security against potentially misaligned AI agents, including threat modeling, capability‑based mitigation ti…
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
Chloe Li, Mary Phuong, Noah Y. Siegel
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38