activity
20162022
most citedAlign before Fuse: Vision and Language Representation Learning with Momentum Distillation

823 citations · 1.2k across the 7 of their papers we have counts for

collaborators

14 papers

cs.CV2022★ 7 cited

TAG: Boosting Text-VQA via Text-aware Visual Question-answer Generation

Jun Wang, Mingfei Gao, Yuqian Hu +5

Text-VQA aims at answering questions that require understanding the textual cues in an image. Despite the great progress of existing Text-VQA methods, their performance suffers fro…

cs.CV2022

Can domain adaptation make object recognition work for everyone?

Viraj Prabhu, Ramprasaath R. Selvaraju, Judy Hoffman +1

Despite the rapid progress in deep visual recognition, modern computer vision datasets significantly overrepresent the developed world and models trained on such datasets underperf…

cs.CV2021

CLIP-Lite: Information Efficient Visual Representation Learning with Language Supervision

Aman Shrivastava, Ramprasaath R. Selvaraju, Nikhil Naik +1

We propose CLIP-Lite, an information efficient method for visual representation learning by feature alignment with textual annotations. Compared to the previously proposed CLIP mod…

cs.CV2021

PreViTS: Contrastive Pretraining with Video Tracking Supervision

Brian Chen, Ramprasaath R. Selvaraju, Shih-Fu Chang +2

Videos are a rich source for self-supervised learning (SSL) of visual representations due to the presence of natural temporal transformations of objects. However, current methods t…

cs.CV2021★ 823 cited

Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare +3

Large-scale vision and language representation learning has shown promising improvements on various vision-language tasks. Most existing methods employ a transformer-based multimod…

cs.CV2020

CASTing Your Model: Learning to Localize Improves Self-Supervised Representations

Ramprasaath R. Selvaraju, Karan Desai, Justin Johnson +1

Recent advances in self-supervised learning (SSL) have largely closed the gap with supervised ImageNet pretraining. Despite their success these methods have been primarily applied…