209 citations · 517 across the 37 of their papers we have counts for
4 papers · 2 filters
Are we asking the right questions in MovieQA?
Bhavan Jasani, Rohit Girdhar, Deva Ramanan
Joint vision and language tasks like visual question answering are fascinating because they explore high-level understanding, but at the same time, can be more prone to language bi…
CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
Rohit Girdhar, Deva Ramanan
Computer vision has undergone a dramatic revolution in performance, driven in large part through deep features trained on large-scale supervised datasets. However, much of these im…
MetaPix: Few-Shot Video Retargeting
Jessica Lee, Deva Ramanan, Rohit Girdhar
We address the task of unsupervised retargeting of human actions from one video to another. We consider the challenging setting where only a few frames of the target is available.…
DistInit: Learning Video Representations Without a Single Labeled Video
Rohit Girdhar, Du Tran, Lorenzo Torresani +1
Video recognition models have progressed significantly over the past few years, evolving from shallow classifiers trained on hand-crafted features to deep spatiotemporal networks.…