44 citations · 92 across the 4 of their papers we have counts for
12 papers
Unsupervised Vision-and-Language Pre-training via Retrieval-based Multi-Granular Alignment
Mingyang Zhou, Licheng Yu, Amanpreet Singh +3
Vision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require p…
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
Douwe Kiela, Hamed Firooz, Aravind Mohan +4
This work proposes a new challenge set for multimodal classification, focusing on detecting hate speech in multimodal memes. It is constructed such that unimodal models struggle an…
Are we pretraining it right? Digging deeper into visio-linguistic pretraining
Amanpreet Singh, Vedanuj Goswami, Devi Parikh
Numerous recent works have proposed pretraining generic visio-linguistic representations and then finetuning them for downstream vision and language tasks. While architecture and o…
TextCaps: a Dataset for Image Captioning with Reading Comprehension
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach +1
Image descriptions can help visually impaired people to quickly understand the image content. While we made significant progress in automatically describing images and optical char…
Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA
Ronghang Hu, Amanpreet Singh, Trevor Darrell +1
Many visual scenes contain text that carries crucial information, and it is thus essential to understand text in images for downstream reasoning tasks. For example, a deep water la…
Towards VQA Models That Can Read
Amanpreet Singh, Vivek Natarajan, Meet Shah +5
Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models ca…