17 citations · 17 across the 1 of their papers we have counts for
1 paper
Dongxu Li, Junnan Li, Hongdong Li +2
Video-and-language pre-training has shown promising improvements on various downstream tasks. Most previous methods capture cross-modal interactions with a transformer-based multim…