3 papers
cs.CR2025
BashArena: A Control Setting for Highly Privileged AI Agents
Adam Kaufman, James Lucassen, Tyler Tracy +2
Future AI agents might run autonomously with elevated privileges. If these agents are misaligned, they might abuse these privileges to cause serious damage. The field of AI control…
cs.CR2025
Factor(T,U): Factored Cognition Strengthens Monitoring of Untrusted AI
Aaron Sandoval, Cody Rushing
The field of AI Control seeks to develop robust control protocols, deployment safeguards for untrusted AI which may be intentionally subversive. However, existing protocols that re…
cs.LG2025
Ctrl-Z: Controlling AI Agents via Resampling
Aryan Bhatt, Cody Rushing, Adam Kaufman +5
Control evaluations measure whether monitoring and security protocols for AI systems prevent intentionally subversive AI models from causing harm. Our work presents the first contr…