162 citations · 166 across the 3 of their papers we have counts for
5 papers · 1 filter
Hopper: Multi-hop Transformer for Spatiotemporal Reasoning
Honglu Zhou, Asim Kadav, Farley Lai +4
This paper considers the problem of spatiotemporal object-centric reasoning in videos. Central to our approach is the notion of object permanence, i.e., the ability to reason about…
15 Keypoints Is All You Need
Michael Snower, Asim Kadav, Farley Lai +1
Pose tracking is an important problem that requires identifying unique human pose-instances and matching them temporally across different frames of a video. However, existing pose…
Contextual Grounding of Natural Language Entities in Images
Farley Lai, Ning Xie, Derek Doran +1
In this paper, we introduce a contextual grounding approach that captures the context in corresponding text entities and image regions to improve the grounding accuracy. Specifical…
Visual Entailment: A Novel Task for Fine-Grained Image Understanding
Ning Xie, Farley Lai, Derek Doran +1
Existing visual reasoning datasets such as Visual Question Answering (VQA), often suffer from biases conditioned on the question, image or answer distributions. The recently propos…
Visual Entailment Task for Visually-Grounded Language Learning
Ning Xie, Farley Lai, Derek Doran +1
We introduce a new inference task - Visual Entailment (VE) - which differs from traditional Textual Entailment (TE) tasks whereby a premise is defined by an image, rather than a na…