1 paper
Zheng Ma, Shi Zong, Mianzhi Pan +4
In recent years, vision and language pre-training (VLP) models have advanced the state-of-the-art results in a variety of cross-modal downstream tasks. Aligning cross-modal semanti…