24 citations · 44 across the 13 of their papers we have counts for
9 papers · 1 filter
Learning Action-Effect Dynamics for Hypothetical Vision-Language Reasoning Task
Shailaja Keyur Sampat, Pratyay Banerjee, Yezhou Yang +1
'Actions' play a vital role in how humans interact with the world. Thus, autonomous agents that would assist us in everyday tasks also require the capability to perform 'Reasoning…
Learning Action-Effect Dynamics from Pairs of Scene-graphs
Shailaja Keyur Sampat, Pratyay Banerjee, Yezhou Yang +1
'Actions' play a vital role in how humans interact with the world. Thus, autonomous agents that would assist us in everyday tasks also require the capability to perform 'Reasoning…
To Find Waldo You Need Contextual Cues: Debiasing Who's Waldo
Yiran Luo, Pratyay Banerjee, Tejas Gokhale +2
We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who's Waldo dataset. Given an image and a caption, PCVG…
Semantically Distributed Robust Optimization for Vision-and-Language Inference
Tejas Gokhale, Abhishek Chaudhary, Pratyay Banerjee +2
Analysis of vision-and-language models has revealed their brittleness under linguistic phenomena such as paraphrasing, negation, textual entailment, and word substitutions with syn…
Weakly Supervised Relative Spatial Reasoning for Visual Question Answering
Pratyay Banerjee, Tejas Gokhale, Yezhou Yang +1
Vision-and-language (V\&L) reasoning necessitates perception of visual concepts such as objects and actions, understanding semantics and language grounding, and reasoning about the…
WeaQA: Weak Supervision via Captions for Visual Question Answering
Pratyay Banerjee, Tejas Gokhale, Yezhou Yang +1
Methodologies for training visual question answering (VQA) models assume the availability of datasets with human-annotated \textit{Image-Question-Answer} (I-Q-A) triplets. This has…