162 citations · 181 across the 6 of their papers we have counts for
11 papers
Hopper: Multi-hop Transformer for Spatiotemporal Reasoning
Honglu Zhou, Asim Kadav, Farley Lai +4
This paper considers the problem of spatiotemporal object-centric reasoning in videos. Central to our approach is the notion of object permanence, i.e., the ability to reason about…
S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation
Yizhe Zhu, Martin Renqiang Min, Asim Kadav +1
We propose a sequential variational autoencoder to learn disentangled representations of sequential data (e.g., videos and audios) under self-supervision. Specifically, we exploit…
15 Keypoints Is All You Need
Michael Snower, Asim Kadav, Farley Lai +1
Pose tracking is an important problem that requires identifying unique human pose-instances and matching them temporally across different frames of a video. However, existing pose…
Contextual Grounding of Natural Language Entities in Images
Farley Lai, Ning Xie, Derek Doran +1
In this paper, we introduce a contextual grounding approach that captures the context in corresponding text entities and image regions to improve the grounding accuracy. Specifical…
Tripping through time: Efficient Localization of Activities in Videos
Meera Hahn, Asim Kadav, James M. Rehg +1
Localizing moments in untrimmed videos via language queries is a new and interesting task that requires the ability to accurately ground language into video. Previous works have ap…
Visual Entailment: A Novel Task for Fine-Grained Image Understanding
Ning Xie, Farley Lai, Derek Doran +1
Existing visual reasoning datasets such as Visual Question Answering (VQA), often suffer from biases conditioned on the question, image or answer distributions. The recently propos…