28 citations · 31 across the 3 of their papers we have counts for
1 paper · 1 filter
Wei Li, Can Gao, Guocheng Niu +5
Vision-Language Pre-training (VLP) has achieved impressive performance on various cross-modal downstream tasks. However, most existing methods can only learn from aligned image-cap…