156 citations · 156 across the 2 of their papers we have counts for
3 papers
cs.CL2021
GEM: A General Evaluation Benchmark for Multimodal Tasks
Lin Su, Nan Duan, Edward Cui +7
In this paper, we present GEM as a General Evaluation benchmark for Multimodal tasks. Different from existing datasets such as GLUE, SuperGLUE, XGLUE and XTREME that mainly focus o…
cs.CV2020★ 156 cited
ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data
Di Qi, Lin Su, Jia Song +3
In this paper, we introduce a new vision-language pre-trained model -- ImageBERT -- for image-text joint embedding. Our model is a Transformer-based model, which takes different mo…
cs.CV2018
Web-Scale Responsive Visual Search at Bing
Houdong Hu, Yan Wang, Linjun Yang +7
In this paper, we introduce a web-scale general visual search system deployed in Microsoft Bing. The system accommodates tens of billions of images in the index, with thousands of…