2 papers
cs.AI2026
Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios
Linus Folkerts, Will Payne, Simon Inman +11
We evaluate the autonomous cyber-attack capabilities of frontier AI models on two purpose-built cyber ranges-a 32-step corporate network attack and a 7-step industrial control syst…
cs.AI2026
Improving Methodologies for LLM Evaluations Across Global Languages
Akriti Vij, Benjamin Chua, Darshini Ramiah +43
As frontier AI models are deployed globally, it is essential that their behaviour remains safe and reliable across diverse linguistic and cultural contexts. To examine how current…