2 papers
cs.CR2026
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks
Michael Kouremetis, Ads Dawson, Raja Sekhar Rao Dheekonda +1
Large language model (LLM) agents routinely cheat on cybersecurity benchmarks, inflating reported pass rates far beyond genuine capability. Prior audits of Cybench found cheating i…
cs.CR2025
The Automation Advantage in AI Red Teaming
Rob Mulla, Ads Dawson, Vincent Abruzzon +4
This paper analyzes Large Language Model (LLM) security vulnerabilities based on data from Crucible, encompassing 214,271 attack attempts by 1,674 users across 30 LLM challenges. O…