156 citations · 243 across the 4 of their papers we have counts for
1 paper · 1 filter
Di Qi, Lin Su, Jia Song +3
In this paper, we introduce a new vision-language pre-trained model -- ImageBERT -- for image-text joint embedding. Our model is a Transformer-based model, which takes different mo…