activity
20172021
most citedS3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation

8 citations · 14 across the 4 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20211 cited

Hopper: Multi-hop Transformer for Spatiotemporal Reasoning

Honglu Zhou, Asim Kadav, Farley Lai +4

This paper considers the problem of spatiotemporal object-centric reasoning in videos. Central to our approach is the notion of object permanence, i.e., the ability to reason about…

cs.CV20208 cited

S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation

Yizhe Zhu, Martin Renqiang Min, Asim Kadav +1

We propose a sequential variational autoencoder to learn disentangled representations of sequential data (e.g., videos and audios) under self-supervision. Specifically, we exploit…

cs.CV2019

15 Keypoints Is All You Need

Michael Snower, Asim Kadav, Farley Lai +1

Pose tracking is an important problem that requires identifying unique human pose-instances and matching them temporally across different frames of a video. However, existing pose…

cs.CV2019

Tripping through time: Efficient Localization of Activities in Videos

Meera Hahn, Asim Kadav, James M. Rehg +1

Localizing moments in untrimmed videos via language queries is a new and interesting task that requires the ability to accurately ground language into video. Previous works have ap…

cs.CV20175 cited

Grounded Objects and Interactions for Video Captioning

Chih-Yao Ma, Asim Kadav, Iain Melvin +3

We address the problem of video captioning by grounding language generation on object interactions in the video. Existing work mostly focuses on overall scene understanding with of…