6 papers
Realistic honeypot evaluations for scheming propensity
Victoria Krakovna, David Lindner, Lewis Ho +2
We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity. Our scheming honeypot evaluations take t…
Gram: Assessing sabotage propensities via automated alignment auditing
David Lindner, Victoria Krakovna, Sebastian Farquhar
We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models across 17 simulated agentic depl…
Evaluating and Understanding Scheming Propensity in LLM Agents
Mia Hopman, Jannes Elstner, Maria Avramidou +2
As frontier language models are increasingly deployed as autonomous agents pursuing complex, long-term objectives, there is increased risk of scheming: agents covertly pursuing mis…
Frontier Models Can Take Actions at Low Probabilities
Alex Serrano, Wen Xing, David Lindner +1
Pre-deployment evaluations inspect only a limited sample of model actions. A malicious model seeking to evade oversight could exploit this by randomizing when to "defect": misbehav…
Practical challenges of control monitoring in frontier AI deployments
David Lindner, Charlie Griffin, Tomek Korbak +4
Automated control monitors could play an important role in overseeing highly capable AI agents that we do not fully trust. Prior work has explored control monitoring in simplified…
Evaluating Frontier Models for Stealth and Situational Awareness
Mary Phuong, Roland S. Zimmermann, Ziyue Wang +6
Recent work has demonstrated the plausibility of frontier AI models scheming -- knowingly and covertly pursuing an objective misaligned with its developer's intentions. Such behavi…