activity
20172022
most citedAttentional Pooling for Action Recognition

209 citations · 241 across the 4 of their papers we have counts for

collaborators

16 papers

cs.CV20222 cited

Learning Video Representations from Large Language Models

Yue Zhao, Ishan Misra, Philipp Krähenbühl +1

We introduce LaViLa, a new approach to learning video-language representations by leveraging Large Language Models (LLMs). We repurpose pre-trained LLMs to be conditioned on visual…

cs.CV2021

An End-to-End Transformer Model for 3D Object Detection

Ishan Misra, Rohit Girdhar, Armand Joulin

We propose 3DETR, an end-to-end Transformer based object detection model for 3D point clouds. Compared to existing detection methods that employ a number of 3D-specific inductive b…

cs.CV2021

Anticipative Video Transformer

Rohit Girdhar, Kristen Grauman

We propose Anticipative Video Transformer (AVT), an end-to-end attention-based video modeling architecture that attends to the previously observed video in order to anticipate futu…

cs.CV2021

3D Spatial Recognition without Spatially Labeled 3D

Zhongzheng Ren, Ishan Misra, Alexander G. Schwing +1

We introduce WyPR, a Weakly-supervised framework for Point cloud Recognition, requiring only scene-level class tags as supervision. WyPR jointly addresses three core 3D recognition…

cs.AI2021

Physical Reasoning Using Dynamics-Aware Models

Eltayeb Ahmed, Anton Bakhtin, Laurens van der Maaten +1

A common approach to solving physical reasoning tasks is to train a value learner on example tasks. A limitation of such an approach is that it requires learning about object dynam…

cs.CV20214 cited

Self-Supervised Pretraining of 3D Features on any Point-Cloud

Zaiwei Zhang, Rohit Girdhar, Armand Joulin +1

Pretraining on large labeled datasets is a prerequisite to achieve good performance in many computer vision tasks like 2D object recognition, video classification etc. However, pre…