activity
20242026
collaborators

11 papers

cs.AI2026

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

Agatha Duzan, Asa Cooper Stickland

Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicit-influence setti…

cs.AI2026

Distributed Attacks in Persistent-State AI Control

Josh Hills, Ida Caspary, Asa Cooper Stickland

As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a…

cs.AI2026

Forecasting Future Behavior as a Learning Task

Mosh Levy, Yoav Goldberg, Asa Cooper Stickland

Trust in an AI system is often anchored by explanations of how it works, which one then uses to forecast its behavior on new inputs. For large reasoning models (LRMs), this convent…

cs.LG2026

Why Do Language Model Agents Whistleblow?

Kushal Agrawal, Frank Xiao, Guido Bergman +1

The deployment of Large Language Models (LLMs) as tool-using agents causes their alignment training to manifest in new ways. Recent work finds that language models can use tools in…

cs.LG2025

Async Control: Stress-testing Asynchronous Control Measures for LLM Agents

Asa Cooper Stickland, Jan Michelfeit, Arathi Mani +6

LLM-based software engineering agents are increasingly used in real-world development tasks, often with access to sensitive data or security-critical codebases. Such agents could i…

cs.LG2025

Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Abhay Sheshadri, Aidan Ewart, Phillip Guo +8

Large language models (LLMs) can often be made to behave in undesirable ways that they are explicitly fine-tuned not to. For example, the LLM red-teaming literature has produced a…