11 papers
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
Agatha Duzan, Asa Cooper Stickland
Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicit-influence setti…
Distributed Attacks in Persistent-State AI Control
Josh Hills, Ida Caspary, Asa Cooper Stickland
As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a…
Forecasting Future Behavior as a Learning Task
Mosh Levy, Yoav Goldberg, Asa Cooper Stickland
Trust in an AI system is often anchored by explanations of how it works, which one then uses to forecast its behavior on new inputs. For large reasoning models (LRMs), this convent…
Why Do Language Model Agents Whistleblow?
Kushal Agrawal, Frank Xiao, Guido Bergman +1
The deployment of Large Language Models (LLMs) as tool-using agents causes their alignment training to manifest in new ways. Recent work finds that language models can use tools in…
Async Control: Stress-testing Asynchronous Control Measures for LLM Agents
Asa Cooper Stickland, Jan Michelfeit, Arathi Mani +6
LLM-based software engineering agents are increasingly used in real-world development tasks, often with access to sensitive data or security-critical codebases. Such agents could i…
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Abhay Sheshadri, Aidan Ewart, Phillip Guo +8
Large language models (LLMs) can often be made to behave in undesirable ways that they are explicitly fine-tuned not to. For example, the LLM red-teaming literature has produced a…