activity
20172021
most citedVisual Entailment: A Novel Task for Fine-Grained Image Understanding

162 citations · 181 across the 6 of their papers we have counts for

collaborators

11 papers

cs.CV20211 cited

Hopper: Multi-hop Transformer for Spatiotemporal Reasoning

Honglu Zhou, Asim Kadav, Farley Lai +4

This paper considers the problem of spatiotemporal object-centric reasoning in videos. Central to our approach is the notion of object permanence, i.e., the ability to reason about…

cs.CV20208 cited

S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation

Yizhe Zhu, Martin Renqiang Min, Asim Kadav +1

We propose a sequential variational autoencoder to learn disentangled representations of sequential data (e.g., videos and audios) under self-supervision. Specifically, we exploit…

cs.CV2019

15 Keypoints Is All You Need

Michael Snower, Asim Kadav, Farley Lai +1

Pose tracking is an important problem that requires identifying unique human pose-instances and matching them temporally across different frames of a video. However, existing pose…

cs.CV20193 cited

Contextual Grounding of Natural Language Entities in Images

Farley Lai, Ning Xie, Derek Doran +1

In this paper, we introduce a contextual grounding approach that captures the context in corresponding text entities and image regions to improve the grounding accuracy. Specifical…

cs.CV2019

Tripping through time: Efficient Localization of Activities in Videos

Meera Hahn, Asim Kadav, James M. Rehg +1

Localizing moments in untrimmed videos via language queries is a new and interesting task that requires the ability to accurately ground language into video. Previous works have ap…

cs.CV2019162 cited

Visual Entailment: A Novel Task for Fine-Grained Image Understanding

Ning Xie, Farley Lai, Derek Doran +1

Existing visual reasoning datasets such as Visual Question Answering (VQA), often suffer from biases conditioned on the question, image or answer distributions. The recently propos…