2 papers
cs.AI2026
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
Justin W. Lin, Eliot Krzysztof Jones, Donovan Julian Jasper +10
We present the first comprehensive evaluation of AI agents against human cybersecurity professionals in a live enterprise environment. We evaluate ten cybersecurity professionals a…
cs.CR2025
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Andy K. Zhang, Neil Perry, Riya Dulepet +24
Language Model (LM) agents for cybersecurity that are capable of autonomously identifying vulnerabilities and executing exploits have potential to cause real-world impact. Policyma…