4 papers
The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning
Wencheng Ye, Yi Bin, Yujuan Ding +7
Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening eviden…
Practical No-box Adversarial Attacks with Training-free Hybrid Image Transformation
Qilong Zhang, Youheng Sun, Chaoning Zhang +4
In recent years, the adversarial vulnerability of deep neural networks (DNNs) has raised increasing attention. Among all the threat models, no-box attacks are the most practical bu…
Informative Scene Graph Generation via Debiasing
Lianli Gao, Xinyu Lyu, Yuyu Guo +5
Scene graph generation aims to detect visual relationship triplets, (subject, predicate, object). Due to biases in data, current models tend to predict common predicates, e.g. "on"…
BadCM: Invisible Backdoor Attack Against Cross-Modal Learning
Zheng Zhang, Xu Yuan, Lei Zhu +2
Despite remarkable successes in unimodal learning tasks, backdoor attacks against cross-modal learning are still underexplored due to the limited generalization and inferior stealt…