3 papers
cs.CR2026
Injection-Execution Dissociation: A Mechanistic Evaluation of Persistent Memory Attacks and Defenses in Stateful LLM Agents
Jun Wen Leong
We discover that prompt-injection success and tool-execution success are separable safety properties: defenses that block injection do not necessarily block execution, and vice ver…
cs.CR2026
Forensic Trajectory Signatures for Agent Memory Poisoning Detection
Jun Wen Leong
We discover a behavioral invariant in LLM agents under persistent memory poisoning and characterize its deployment boundary. In architectures where retrieval is routed through obse…
cs.LG2026
Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers
Jun Wen Leong
Reasoning models deployed as safety monitors exhibit a systematic vulnerability: reasoning-token budget starvation. Adversarial inputs require more reasoning tokens tha…