27 citations · 63 across the 7 of their papers we have counts for
1 paper · 1 filter
Liunian Harold Li, Mark Yatskar, Da Yin +2
We propose VisualBERT, a simple and flexible framework for modeling a broad range of vision-and-language tasks. VisualBERT consists of a stack of Transformer layers that implicitly…