1.4k citations · 2.9k across the 107 of their papers we have counts for
1 paper · 2 filters
Liunian Harold Li, Mark Yatskar, Da Yin +2
We propose VisualBERT, a simple and flexible framework for modeling a broad range of vision-and-language tasks. VisualBERT consists of a stack of Transformer layers that implicitly…