activity
20182025
most citedLearning when to Communicate at Scale in Multiagent Cooperative and Competitive Tasks

44 citations · 92 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2025

LLaVA-RE: Binary Image-Text Relevancy Evaluation with Multimodal Large Language Model

Tao Sun, Oliver Liu, JinJin Li +1

Multimodal generative AI usually involves generating image or text responses given inputs in another modality. The evaluation of image-text relevancy is essential for measuring res…

cs.CV2025

Improved Alignment of Modalities in Large Vision Language Models

Kartik Jangra, Aman Kumar Singh, Yashwani Mann +1

Recent advancements in vision-language models have achieved remarkable results in making language models understand vision inputs. However, a unified approach to align these models…

cs.CV2022★ 1 cited

Unsupervised Vision-and-Language Pre-training via Retrieval-based Multi-Granular Alignment

Mingyang Zhou, Licheng Yu, Amanpreet Singh +3

Vision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require p…

cs.CV2020★ 30 cited

Are we pretraining it right? Digging deeper into visio-linguistic pretraining

Amanpreet Singh, Vedanuj Goswami, Devi Parikh

Numerous recent works have proposed pretraining generic visio-linguistic representations and then finetuning them for downstream vision and language tasks. While architecture and o…

cs.CV2020

TextCaps: a Dataset for Image Captioning with Reading Comprehension

Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach +1

Image descriptions can help visually impaired people to quickly understand the image content. While we made significant progress in automatically describing images and optical char…

cs.CV2019

Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA

Ronghang Hu, Amanpreet Singh, Trevor Darrell +1

Many visual scenes contain text that carries crucial information, and it is thus essential to understand text in images for downstream reasoning tasks. For example, a deep water la…