13 citations · 14 across the 10 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models
Chengyin Hu, Xuemeng Sun, Jiaju Han +5
Visual-Language Models (VLMs) have demonstrated exceptional cross-modal understanding across various tasks, including zero-shot classification, image captioning, and visual questio…
cs.CV2022★ 1 cited
Aligning Source Visual and Target Language Domains for Unpaired Video Captioning
Fenglin Liu, Xian Wu, Chenyu You +3
Training supervised video captioning model requires coupled video-caption pairs. However, for many targeted languages, sufficient paired data are not available. To this end, we int…