3 papers
cs.CR2026
Amplify, Don't Create: Temporal Accumulation for Slow-Burn Prompt Injection
J Alex Corll
Most prompt-injection detectors score a single event or message. Control-plane attacks against tool-using agents can instead distribute weak directives across a trajectory while ke…
cs.CR2026
The Mirror Design Pattern: Strict Data Geometry over Model Scale for Prompt Injection Detection
J Alex Corll
Prompt injection defenses are often framed as semantic understanding problems and delegated to increasingly large neural detectors. For the first screening layer, however, the requ…
cs.CR2026
Peak + Accumulation: A Proxy-Level Scoring Formula for Multi-Turn LLM Attack Detection
J Alex Corll
Multi-turn prompt injection attacks distribute malicious intent across multiple conversation turns, exploiting the assumption that each turn is evaluated independently. While singl…