209 citations · 241 across the 4 of their papers we have counts for
16 papers
Learning Video Representations from Large Language Models
Yue Zhao, Ishan Misra, Philipp Krähenbühl +1
We introduce LaViLa, a new approach to learning video-language representations by leveraging Large Language Models (LLMs). We repurpose pre-trained LLMs to be conditioned on visual…
An End-to-End Transformer Model for 3D Object Detection
Ishan Misra, Rohit Girdhar, Armand Joulin
We propose 3DETR, an end-to-end Transformer based object detection model for 3D point clouds. Compared to existing detection methods that employ a number of 3D-specific inductive b…
Anticipative Video Transformer
Rohit Girdhar, Kristen Grauman
We propose Anticipative Video Transformer (AVT), an end-to-end attention-based video modeling architecture that attends to the previously observed video in order to anticipate futu…
3D Spatial Recognition without Spatially Labeled 3D
Zhongzheng Ren, Ishan Misra, Alexander G. Schwing +1
We introduce WyPR, a Weakly-supervised framework for Point cloud Recognition, requiring only scene-level class tags as supervision. WyPR jointly addresses three core 3D recognition…
Physical Reasoning Using Dynamics-Aware Models
Eltayeb Ahmed, Anton Bakhtin, Laurens van der Maaten +1
A common approach to solving physical reasoning tasks is to train a value learner on example tasks. A limitation of such an approach is that it requires learning about object dynam…
Self-Supervised Pretraining of 3D Features on any Point-Cloud
Zaiwei Zhang, Rohit Girdhar, Armand Joulin +1
Pretraining on large labeled datasets is a prerequisite to achieve good performance in many computer vision tasks like 2D object recognition, video classification etc. However, pre…