4 papers
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
Justin W. Lin, Eliot Krzysztof Jones, Donovan Julian Jasper +10
We present the first comprehensive evaluation of AI agents against human cybersecurity professionals in a live enterprise environment. We evaluate ten cybersecurity professionals a…
Cryptographic Data Exchange for Nuclear Warheads
Neil Perry, Daniil Zhukov
Nuclear arms control treaties have historically focused on strategic nuclear delivery systems, indirectly restricting strategic nuclear warhead numbers and leaving nonstrategic nuc…
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Andy K. Zhang, Neil Perry, Riya Dulepet +24
Language Model (LM) agents for cybersecurity that are capable of autonomously identifying vulnerabilities and executing exploits have potential to cause real-world impact. Policyma…
Robust Steganography from Large Language Models
Neil Perry, Sanket Gupte, Nishant Pitta +1
Recent steganographic schemes, starting with Meteor (CCS'21), rely on leveraging large language models (LLMs) to resolve a historically-challenging task of disguising covert commun…