works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.CR2026

GDM AI Control Roadmap

Mary Phuong, Erik Jenner, Laurent Simon +4

The paper presents the GDM AI Control Roadmap, a framework for internal security against potentially misaligned AI agents, including threat modeling, capability‑based mitigation ti…

cs.LG2026

Realistic honeypot evaluations for scheming propensity

Victoria Krakovna, David Lindner, Lewis Ho +2

We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity. Our scheming honeypot evaluations take t…

cs.LG2026

Gram: Assessing sabotage propensities via automated alignment auditing

David Lindner, Victoria Krakovna, Sebastian Farquhar

We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models across 17 simulated agentic depl…

cs.LG2026

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs

Eric Easley, Sebastian Farquhar

We address jailbreaks, backdoors, and unlearning for large language models (LLMs). Unlike prior work, which trains LLMs based on their actions when given malign instructions, our m…

cs.CR2025

Practical challenges of control monitoring in frontier AI deployments

David Lindner, Charlie Griffin, Tomek Korbak +4

Automated control monitors could play an important role in overseeing highly capable AI agents that we do not fully trust. Prior work has explored control monitoring in simplified…

cs.LG2025

MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking

Sebastian Farquhar, Vikrant Varma, David Lindner +4

Future advanced AI systems may learn sophisticated strategies through reinforcement learning (RL) that humans cannot understand well enough to safely evaluate. We propose a trainin…