most citedImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

156 citations · 243 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL2021

GEM: A General Evaluation Benchmark for Multimodal Tasks

Lin Su, Nan Duan, Edward Cui +7

In this paper, we present GEM as a General Evaluation benchmark for Multimodal tasks. Different from existing datasets such as GLUE, SuperGLUE, XGLUE and XTREME that mainly focus o…

cs.CL2020

M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-training

Minheng Ni, Haoyang Huang, Lin Su +6

We present M3P, a Multitask Multilingual Multimodal Pre-trained model that combines multilingual pre-training and multimodal pre-training into a unified framework via multitask pre…

cs.CL202067 cited

XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation

Yaobo Liang, Nan Duan, Yeyun Gong +21

In this paper, we introduce XGLUE, a new benchmark dataset that can be used to train large-scale cross-lingual pre-trained models using multilingual and bilingual corpora and evalu…

cs.CL202020 cited

XGPT: Cross-modal Generative Pre-Training for Image Captioning

Qiaolin Xia, Haoyang Huang, Nan Duan +7

While many BERT-based cross-modal pre-trained models produce excellent results on downstream understanding tasks like image-text retrieval and VQA, they cannot be applied to genera…

cs.CV2020156 cited

ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

Di Qi, Lin Su, Jia Song +3

In this paper, we introduce a new vision-language pre-trained model -- ImageBERT -- for image-text joint embedding. Our model is a Transformer-based model, which takes different mo…