1 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.AI2026★ 1 cited
Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios
Linus Folkerts, Will Payne, Simon Inman +11
We evaluate the autonomous cyber-attack capabilities of frontier AI models on two purpose-built cyber ranges-a 32-step corporate network attack and a 7-step industrial control syst…
cs.CR2026★ 1 cited
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
Rahul Marchand, Art O Cathain, Jerome Wynne +8
Large language models (LLMs) increasingly act as autonomous agents, using tools to execute code, read and write files, and access networks, creating novel security risks. To mitiga…