3 papers
cs.CR2026
Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models
Xinwei Zhang, Li Bai, Tianwei Zhang +5
Large vision-language models (LVLMs) have achieved impressive performance across multimodal tasks, but their reliance on visual inputs exposes them to adversarial threats. Encoder-…
cs.CR2026
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
Xinwei Zhang, Hangcheng Liu, Li Bai +4
Visual token compression is widely used to accelerate large vision-language models (LVLMs) by pruning or merging visual tokens, yet its adversarial robustness remains unexplored. W…
cs.CV2026
Physical Prompt Injection Attacks on Large Vision-Language Models
Chen Ling, Kai Hu, Hangcheng Liu +3
Large Vision-Language Models (LVLMs) are increasingly deployed in real-world intelligent systems for perception and reasoning in open physical environments. While LVLMs are known t…