Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Steal the Patch Size: Adversarially Manipulate Vision-Language Models
Kai Hu, Akash Bharadwaj, Weichen Yu +1
We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including the visual patch size and in…
cs.CV2025
Transferable Adversarial Attacks on Black-Box Vision-Language Models
Kai Hu, Weichen Yu, Li Zhang +5
Vision Large Language Models (VLLMs) are increasingly deployed to offer advanced capabilities on inputs comprising both text and images. While prior research has shown that adversa…
cs.CV2024
Is Your Text-to-Image Model Robust to Caption Noise?
Weichen Yu, Ziyan Yang, Shanchuan Lin +5
In text-to-image (T2I) generation, a prevalent training technique involves utilizing Vision Language Models (VLMs) for image re-captioning. Even though VLMs are known to exhibit ha…