8 citations · 14 across the 4 of their papers we have counts for
5 papers · 1 filter
Hopper: Multi-hop Transformer for Spatiotemporal Reasoning
Honglu Zhou, Asim Kadav, Farley Lai +4
This paper considers the problem of spatiotemporal object-centric reasoning in videos. Central to our approach is the notion of object permanence, i.e., the ability to reason about…
S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation
Yizhe Zhu, Martin Renqiang Min, Asim Kadav +1
We propose a sequential variational autoencoder to learn disentangled representations of sequential data (e.g., videos and audios) under self-supervision. Specifically, we exploit…
15 Keypoints Is All You Need
Michael Snower, Asim Kadav, Farley Lai +1
Pose tracking is an important problem that requires identifying unique human pose-instances and matching them temporally across different frames of a video. However, existing pose…
Tripping through time: Efficient Localization of Activities in Videos
Meera Hahn, Asim Kadav, James M. Rehg +1
Localizing moments in untrimmed videos via language queries is a new and interesting task that requires the ability to accurately ground language into video. Previous works have ap…
Grounded Objects and Interactions for Video Captioning
Chih-Yao Ma, Asim Kadav, Iain Melvin +3
We address the problem of video captioning by grounding language generation on object interactions in the video. Existing work mostly focuses on overall scene understanding with of…