2 papers
cs.CR2026
Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents
Harry Owiredu-Ashley
Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary attack-success rate…
cs.CR2026
ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models
Harry Owiredu-Ashley
Most adversarial evaluations of large language model (LLM) safety assess single prompts and report binary pass/fail outcomes, which fails to capture how safety properties evolve un…