24 citations · 37 across the 10 of their papers we have counts for
20 papers
Learning Action-Effect Dynamics for Hypothetical Vision-Language Reasoning Task
Shailaja Keyur Sampat, Pratyay Banerjee, Yezhou Yang +1
'Actions' play a vital role in how humans interact with the world. Thus, autonomous agents that would assist us in everyday tasks also require the capability to perform 'Reasoning…
Learning Action-Effect Dynamics from Pairs of Scene-graphs
Shailaja Keyur Sampat, Pratyay Banerjee, Yezhou Yang +1
'Actions' play a vital role in how humans interact with the world. Thus, autonomous agents that would assist us in everyday tasks also require the capability to perform 'Reasoning…
To Find Waldo You Need Contextual Cues: Debiasing Who's Waldo
Yiran Luo, Pratyay Banerjee, Tejas Gokhale +2
We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who's Waldo dataset. Given an image and a caption, PCVG…
Weakly-Supervised Visual-Retriever-Reader for Knowledge-based Question Answering
Man Luo, Yankai Zeng, Pratyay Banerjee +1
Knowledge-based visual question answering (VQA) requires answering questions with external knowledge in addition to the content of images. One dataset that is mostly used in evalua…
Weakly Supervised Relative Spatial Reasoning for Visual Question Answering
Pratyay Banerjee, Tejas Gokhale, Yezhou Yang +1
Vision-and-language (V\&L) reasoning necessitates perception of visual concepts such as objects and actions, understanding semantics and language grounding, and reasoning about the…
Constructing Flow Graphs from Procedural Cybersecurity Texts
Kuntal Kumar Pal, Kazuaki Kashihara, Pratyay Banerjee +3
Following procedural texts written in natural languages is challenging. We must read the whole text to identify the relevant information or identify the instruction flows to comple…