4 citations · 5 across the 2 of their papers we have counts for
1 paper · 1 filter
Fei Yu, Jiji Tang, Weichong Yin +4
We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL…