74 citations · 132 across the 19 of their papers we have counts for
25 papers
Learning Action-Effect Dynamics for Hypothetical Vision-Language Reasoning Task
Shailaja Keyur Sampat, Pratyay Banerjee, Yezhou Yang +1
'Actions' play a vital role in how humans interact with the world. Thus, autonomous agents that would assist us in everyday tasks also require the capability to perform 'Reasoning…
Learning Action-Effect Dynamics from Pairs of Scene-graphs
Shailaja Keyur Sampat, Pratyay Banerjee, Yezhou Yang +1
'Actions' play a vital role in how humans interact with the world. Thus, autonomous agents that would assist us in everyday tasks also require the capability to perform 'Reasoning…
CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering
Maitreya Patel, Tejas Gokhale, Chitta Baral +1
Videos often capture objects, their visible properties, their motion, and the interactions between different objects. Objects also have physical properties such as mass, which the…
Tragedy Plus Time: Capturing Unintended Human Activities from Weakly-labeled Videos
Arnav Chakravarthy, Zhiyuan Fang, Yezhou Yang
In videos that contain actions performed unintentionally, agents do not achieve their desired goals. In such videos, it is challenging for computer vision systems to understand hig…
SSR-GNNs: Stroke-based Sketch Representation with Graph Neural Networks
Sheng Cheng, Yi Ren, Yezhou Yang
This paper follows cognitive studies to investigate a graph representation for sketches, where the information of strokes, i.e., parts of a sketch, are encoded on vertices and info…
To Find Waldo You Need Contextual Cues: Debiasing Who's Waldo
Yiran Luo, Pratyay Banerjee, Tejas Gokhale +2
We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who's Waldo dataset. Given an image and a caption, PCVG…