4 papers
Code Monitor Red Teaming for Public-Test-Passing Code
Junchi Liao, Jiawen Deng, Fuji Ren
Visible tests are a common gate for LLM-generated code, but passing them does not certify specification correctness. We study a deployment-like monitoring problem: after code has p…
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
Monika JotautaitÄ, Maria Angelica Martinez, Ollie Matthews +1
We introduce a red-teaming methodology that exposes harder-to-catch attacks for coding-agent monitors, suggesting that current practices may under-elicit attacks and overstate moni…
Async Control: Stress-testing Asynchronous Control Measures for LLM Agents
Asa Cooper Stickland, Jan Michelfeit, Arathi Mani +6
LLM-based software engineering agents are increasingly used in real-world development tasks, often with access to sensitive data or security-critical codebases. Such agents could i…
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
Sid Black, Asa Cooper Stickland, Jake Pencharz +7
Uncontrollable autonomous replication of language model agents poses a critical safety risk. To better understand this risk, we introduce RepliBench, a suite of evaluations designe…