Showing cs.CRShow all
2 papers · 1 filter
cs.CR2026
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents
Shane Caldwell, Max Harley, Ads Dawson +3
As LLM agents take on offensive security work, a single out-of-scope tool call can breach a client's engagement boundary, disrupt production, or void a bug-bounty finding. Unlike a…
cs.CR2025
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
Ads Dawson, Rob Mulla, Nick Landers +1
We introduce AIRTBench, an AI red teaming benchmark for evaluating language models' ability to autonomously discover and exploit Artificial Intelligence and Machine Learning (AI/ML…