12 papers
Code Monitor Red Teaming for Public-Test-Passing Code
Junchi Liao, Jiawen Deng, Fuji Ren
Visible tests are a common gate for LLM-generated code, but passing them does not certify specification correctness. We study a deployment-like monitoring problem: after code has p…
Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety
Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad +3
An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework for deploying capable but unt…
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
Elle Najt, Colin Toft, Tyler Tracy +2
Since autonomous coding agents generate complex behaviors at high-volume, we may want to use other LLMs to monitor actions to reduce the risk from dangerous misaligned behavior. To…
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
Monika JotautaitÄ, Maria Angelica Martinez, Ollie Matthews +1
We introduce a red-teaming methodology that exposes harder-to-catch attacks for coding-agent monitors, suggesting that current practices may under-elicit attacks and overstate moni…
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
Tyler Tracy, Ram Potham, Nick Kuhn +31
We introduce LinuxArena, a control setting in which agents operate directly on live, multi-service production environments. LinuxArena contains 20 environments, 1,671 main tasks re…
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
Joachim Schaeffer, Arjun Khandelwal, Tyler Tracy
Future AI deployments will likely be monitored for malicious behaviour. The ability of these AIs to subvert monitors by adversarially selecting against them - attack selection - is…