8 citations · 8 across the 1 of their papers we have counts for
1 paper
Zixin Guo, Tzu-Jui Julius Wang, Selen Pehlivan +2
Vision-language (VL) Pre-training (VLP) has shown to well generalize VL models over a wide range of VL downstream tasks, especially for cross-modal retrieval. However, it hinges on…