2 papers
cs.CR2026
GhostPrompt: Cross-Image Adversarial Prompt for Vision-Language Models
Li Zeng, Zeyu Ye, Meng Xie +4
Vision-Language Models (VLMs) are known to be vulnerable to adversarial attacks, where subtle perturbations to images or texts induce erroneous outputs. However, most text-based at…
cs.CR2026
Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics
Hangtao Zhang, Yucheng Zhao, Sishun Liu +8
Jailbreak prompts can bypass alignment guardrails in large language models (LLMs) and elicit unsafe outputs, making reliable deployment-time detection critical. Prior detection app…