1 paper · 1 filter
Junhui Wang, Hangtao Zhang, Zhirun Zheng +5
The paper introduces PVDetector, a training‑free method that detects prompt injection attacks on purpose‑specific LLM agents by measuring alignment of hidden states with policy‑vio…