activity
20192024
most citedKnowledge Fusion and Semantic Knowledge Ranking for Open Domain Question Answering

24 citations · 44 across the 13 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2022

Learning Action-Effect Dynamics for Hypothetical Vision-Language Reasoning Task

Shailaja Keyur Sampat, Pratyay Banerjee, Yezhou Yang +1

'Actions' play a vital role in how humans interact with the world. Thus, autonomous agents that would assist us in everyday tasks also require the capability to perform 'Reasoning…

cs.CV2022★ 1 cited

Learning Action-Effect Dynamics from Pairs of Scene-graphs

Shailaja Keyur Sampat, Pratyay Banerjee, Yezhou Yang +1

'Actions' play a vital role in how humans interact with the world. Thus, autonomous agents that would assist us in everyday tasks also require the capability to perform 'Reasoning…

cs.CV2022

To Find Waldo You Need Contextual Cues: Debiasing Who's Waldo

Yiran Luo, Pratyay Banerjee, Tejas Gokhale +2

We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who's Waldo dataset. Given an image and a caption, PCVG…

cs.CV2021

Semantically Distributed Robust Optimization for Vision-and-Language Inference

Tejas Gokhale, Abhishek Chaudhary, Pratyay Banerjee +2

Analysis of vision-and-language models has revealed their brittleness under linguistic phenomena such as paraphrasing, negation, textual entailment, and word substitutions with syn…

cs.CV2021★ 1 cited

Weakly Supervised Relative Spatial Reasoning for Visual Question Answering

Pratyay Banerjee, Tejas Gokhale, Yezhou Yang +1

Vision-and-language (V\&L) reasoning necessitates perception of visual concepts such as objects and actions, understanding semantics and language grounding, and reasoning about the…

cs.CV2020

WeaQA: Weak Supervision via Captions for Visual Question Answering

Pratyay Banerjee, Tejas Gokhale, Yezhou Yang +1

Methodologies for training visual question answering (VQA) models assume the availability of datasets with human-annotated \textit{Image-Question-Answer} (I-Q-A) triplets. This has…