1 paper
Jiwei Guan, Tianyu Ding, Longbing Cao +3
Vision-language pretraining (VLP) with transformers has demonstrated exceptional performance across numerous multimodal tasks. However, the adversarial robustness of these models h…