5 papers · 1 filter
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
Elle Najt, Colin Toft, Tyler Tracy +2
Since autonomous coding agents generate complex behaviors at high-volume, we may want to use other LLMs to monitor actions to reduce the risk from dangerous misaligned behavior. To…
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
Monika Jotautaitė, Maria Angelica Martinez, Ollie Matthews +1
We introduce a red-teaming methodology that exposes harder-to-catch attacks for coding-agent monitors, suggesting that current practices may under-elicit attacks and overstate moni…
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
Tyler Tracy, Ram Potham, Nick Kuhn +31
We introduce LinuxArena, a control setting in which agents operate directly on live, multi-service production environments. LinuxArena contains 20 environments, 1,671 main tasks re…
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
Joachim Schaeffer, Arjun Khandelwal, Tyler Tracy
Future AI deployments will likely be monitored for malicious behaviour. The ability of these AIs to subvert monitors by adversarially selecting against them - attack selection - is…
BashArena: A Control Setting for Highly Privileged AI Agents
Adam Kaufman, James Lucassen, Tyler Tracy +2
Future AI agents might run autonomously with elevated privileges. If these agents are misaligned, they might abuse these privileges to cause serious damage. The field of AI control…