3 papers
cs.AI2026
Retrying vs Resampling in AI Control
James Lucassen, Adam Kaufman
AI coding scaffolds like Claude Code and Codex use retrying: blocking actions flagged as risky and continuing the trajectory. We study retrying from an AI control perspective, whic…
cs.CR2025
BashArena: A Control Setting for Highly Privileged AI Agents
Adam Kaufman, James Lucassen, Tyler Tracy +2
Future AI agents might run autonomously with elevated privileges. If these agents are misaligned, they might abuse these privileges to cause serious damage. The field of AI control…
cs.LG2025
Ctrl-Z: Controlling AI Agents via Resampling
Aryan Bhatt, Cody Rushing, Adam Kaufman +5
Control evaluations measure whether monitoring and security protocols for AI systems prevent intentionally subversive AI models from causing harm. Our work presents the first contr…