Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026★ 1 cited
Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios
Linus Folkerts, Will Payne, Simon Inman +11
We evaluate the autonomous cyber-attack capabilities of frontier AI models on two purpose-built cyber ranges-a 32-step corporate network attack and a 7-step industrial control syst…
cs.AI2025
PaperBench: Evaluating AI's Ability to Replicate AI Research
Giulio Starace, Oliver Jaffe, Dane Sherburn +10
We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research. Agents must replicate 20 ICML 2024 Spotlight and Oral papers fro…