195 citations · 354 across the 20 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2022★ 1 cited
CLOP: Video-and-Language Pre-Training with Knowledge Regularizations
Guohao Li, Hu Yang, Feng He +4
Video-and-language pre-training has shown promising results for learning generalizable representations. Most existing approaches usually model video and text in an implicit manner,…
cs.CV2022★ 2 cited
UNIMO-2: End-to-End Unified Vision-Language Grounded Learning
Wei Li, Can Gao, Guocheng Niu +5
Vision-Language Pre-training (VLP) has achieved impressive performance on various cross-modal downstream tasks. However, most existing methods can only learn from aligned image-cap…
cs.CV2020
ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph
Fei Yu, Jiji Tang, Weichong Yin +4
We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL…