3 papers
cs.CR2026
TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models
Meng Xie, Li Zeng, Hangtao Zhang +4
Recent commercial image-generation models can generate high-quality images with readable text (e.g., posters, infographics, and manuals), attracting considerable attention. Yet we…
cs.CR2026
GhostPrompt: Cross-Image Adversarial Prompt for Vision-Language Models
Li Zeng, Zeyu Ye, Meng Xie +4
Vision-Language Models (VLMs) are known to be vulnerable to adversarial attacks, where subtle perturbations to images or texts induce erroneous outputs. However, most text-based at…
cs.CR2026
PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis
Junhui Wang, Hangtao Zhang, Zhirun Zheng +5
The paper introduces PVDetector, a training‑free method that detects prompt injection attacks on purpose‑specific LLM agents by measuring alignment of hidden states with policy‑vio…