4 papers
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks
Michael Kouremetis, Ads Dawson, Raja Sekhar Rao Dheekonda +1
Large language model (LLM) agents routinely cheat on cybersecurity benchmarks, inflating reported pass rates far beyond genuine capability. Prior audits of Cybench found cheating i…
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents
Shane Caldwell, Max Harley, Ads Dawson +3
As LLM agents take on offensive security work, a single out-of-scope tool call can breach a client's engagement boundary, disrupt production, or void a bug-bounty finding. Unlike a…
PentestJudge: Judging Agent Behavior Against Operational Requirements
Shane Caldwell, Max Harley, Michael Kouremetis +2
We introduce PentestJudge, a system for evaluating the operations of penetration testing agents. PentestJudge is a large language model (LLM)-as-judge with access to tools that all…
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
Michael Kouremetis, Marissa Dotter, Alex Byrne +5
The prospect of artificial intelligence (AI) competing in the adversarial landscape of cyber security has long been considered one of the most impactful, challenging, and potential…