44 citations · 92 across the 7 of their papers we have counts for
7 papers · 1 filter
LLaVA-RE: Binary Image-Text Relevancy Evaluation with Multimodal Large Language Model
Tao Sun, Oliver Liu, JinJin Li +1
Multimodal generative AI usually involves generating image or text responses given inputs in another modality. The evaluation of image-text relevancy is essential for measuring res…
Improved Alignment of Modalities in Large Vision Language Models
Kartik Jangra, Aman Kumar Singh, Yashwani Mann +1
Recent advancements in vision-language models have achieved remarkable results in making language models understand vision inputs. However, a unified approach to align these models…
Unsupervised Vision-and-Language Pre-training via Retrieval-based Multi-Granular Alignment
Mingyang Zhou, Licheng Yu, Amanpreet Singh +3
Vision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require p…
Are we pretraining it right? Digging deeper into visio-linguistic pretraining
Amanpreet Singh, Vedanuj Goswami, Devi Parikh
Numerous recent works have proposed pretraining generic visio-linguistic representations and then finetuning them for downstream vision and language tasks. While architecture and o…
TextCaps: a Dataset for Image Captioning with Reading Comprehension
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach +1
Image descriptions can help visually impaired people to quickly understand the image content. While we made significant progress in automatically describing images and optical char…
Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA
Ronghang Hu, Amanpreet Singh, Trevor Darrell +1
Many visual scenes contain text that carries crucial information, and it is thus essential to understand text in images for downstream reasoning tasks. For example, a deep water la…