1 paper
Youze Wang, Wenbo Hu, Yinpeng Dong +3
The integration of visual and textual data in Vision-Language Pre-training (VLP) models is crucial for enhancing vision-language understanding. However, the adversarial robustness…