6 papers
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
Ads Dawson, Adrian Wood
Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectabl…
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks
Michael Kouremetis, Ads Dawson, Raja Sekhar Rao Dheekonda +1
Large language model (LLM) agents routinely cheat on cybersecurity benchmarks, inflating reported pass rates far beyond genuine capability. Prior audits of Cybench found cheating i…
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents
Shane Caldwell, Max Harley, Ads Dawson +3
As LLM agents take on offensive security work, a single out-of-scope tool call can breach a client's engagement boundary, disrupt production, or void a bug-bounty finding. Unlike a…
MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm
Vineeth Sai Narajala, Manish Bhatt, Idan Habler +2
The AI trustworthiness crisis threatens to derail the artificial intelligence revolution, with regulatory barriers, security vulnerabilities, and accountability gaps preventing dep…
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
Ads Dawson, Rob Mulla, Nick Landers +1
We introduce AIRTBench, an AI red teaming benchmark for evaluating language models' ability to autonomously discover and exploit Artificial Intelligence and Machine Learning (AI/ML…
The Automation Advantage in AI Red Teaming
Rob Mulla, Ads Dawson, Vincent Abruzzon +4
This paper analyzes Large Language Model (LLM) security vulnerabilities based on data from Crucible, encompassing 214,271 attack attempts by 1,674 users across 30 LLM challenges. O…