1 citations · 1 across the 16 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
SAVER: Selective As-Needed Vision Evidence for Multimodal Information Extraction
Miaobo Hu, Shuhao Hu, Bokun Wang +5
Multimodal IE in social media is difficult because a post may attach multiple images that are weakly related, redundant, or even misleading with respect to the text. In this settin…
cs.CV2026
TAME: Test-Time Adversarial Prompt Tuning via Mixture-of-Experts for Vision-Language Models
Xin Wang, Yixu Wang, Jiaming Zhang +6
Large-scale pre-trained Vision-Language models (VLMs), such as CLIP, exhibit strong zero-shot generalization, yet remain highly vulnerable to imperceptible adversarial perturbation…