3 citations · 4 across the 4 of their papers we have counts for
1 paper · 1 filter
Dongchen Han, Xiaojun Jia, Yang Bai +3
Vision-language pre-training (VLP) models demonstrate impressive abilities in processing both images and text. However, they are vulnerable to multi-modal adversarial examples (AEs…