3 papers
cs.CR2026
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
Rahul Marchand, Art O Cathain, Jerome Wynne +8
Large language models (LLMs) increasingly act as autonomous agents, using tools to execute code, read and write files, and access networks, creating novel security risks. To mitiga…
cs.AI2026
Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios
Linus Folkerts, Will Payne, Simon Inman +11
We evaluate the autonomous cyber-attack capabilities of frontier AI models on two purpose-built cyber ranges-a 32-step corporate network attack and a 7-step industrial control syst…
cs.AI2026
Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats
Ee Wei Seah, Yongsen Zheng, Naga Nikshith +67
The rapid rise of autonomous AI systems and advancements in agent capabilities are introducing new risks due to reduced oversight of real-world interactions. Yet agent testing rema…