Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours
Raja Sekhar Rao Dheekonda, Will Pearce, Nick Landers
AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a primary defense, current app…
cs.AI2025
PentestJudge: Judging Agent Behavior Against Operational Requirements
Shane Caldwell, Max Harley, Michael Kouremetis +2
We introduce PentestJudge, a system for evaluating the operations of penetration testing agents. PentestJudge is a large language model (LLM)-as-judge with access to tools that all…