3 papers
cs.CR2026
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks
Michael Kouremetis, Ads Dawson, Raja Sekhar Rao Dheekonda +1
Large language model (LLM) agents routinely cheat on cybersecurity benchmarks, inflating reported pass rates far beyond genuine capability. Prior audits of Cybench found cheating i…
cs.AI2026
Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours
Raja Sekhar Rao Dheekonda, Will Pearce, Nick Landers
AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a primary defense, current app…
cs.AI2025
Lessons From Red Teaming 100 Generative AI Products
Blake Bullwinkel, Amanda Minnich, Shiven Chawla +23
In recent years, AI red teaming has emerged as a practice for probing the safety and security of generative AI systems. Due to the nascency of the field, there are many open questi…