4 papers · 1 filter
Readable Yet Unpredictable: Rotated-Outcome Prediction in Vision-Language Models
Lexin Wang, Shenghua Liu, Yiwei Wang +2
Can vision-language models predict what a 180° rotation would reveal from the original image alone? We study this ability through Rotated-Outcome Prediction: given an original ima…
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
Wanqing Cui, Rui Cheng, Jiafeng Guo +1
Existing two-stream models, such as CLIP, encode images and text through independent representations, showing good performance while ensuring retrieval speed, have attracted attent…
Classifier Guidance Enhances Diffusion-based Adversarial Purification by Preserving Predictive Information
Mingkun Zhang, Jianing Li, Wei Chen +2
Adversarial purification is one of the promising approaches to defend neural networks against adversarial attacks. Recently, methods utilizing diffusion probabilistic models have a…
Visual Transformation Telling
Wanqing Cui, Xin Hong, Yanyan Lan +3
Humans can naturally reason from superficial state differences (e.g. ground wetness) to transformations descriptions (e.g. raining) according to their life experience. In this pape…