198 citations · 802 across the 29 of their papers we have counts for
19 papers · 1 filter
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles
Chaitanya Ryali, Yuan-Ting Hu, Daniel Bolya +10
Modern hierarchical vision transformers have added several vision-specific components in the pursuit of supervised classification performance. While these components lead to effect…
Navigating to Objects Specified by Images
Jacob Krantz, Theophile Gervet, Karmesh Yadav +7
Images are a convenient way to specify which particular object instance an embodied agent should navigate to. Solving this task requires semantic visual reasoning and exploration o…
Decoupling Human and Camera Motion from Videos in the Wild
Vickie Ye, Georgios Pavlakos, Jitendra Malik +1
We propose a method to reconstruct global human trajectories from videos in the wild. Our optimization method decouples the camera and human motion, which allows us to place people…
Tracking People by Predicting 3D Appearance, Location & Pose
Jathushan Rajasegaran, Georgios Pavlakos, Angjoo Kanazawa +1
In this paper, we present an approach for tracking people in monocular videos, by predicting their future 3D representations. To achieve this, we first lift people to 3D from a sin…
SEAL: Self-supervised Embodied Active Learning using Exploration and 3D Consistency
Devendra Singh Chaplot, Murtaza Dalal, Saurabh Gupta +2
In this paper, we explore how we can build upon the data and models of Internet images and use them to adapt to robot vision without requiring any extra labels. We present a framew…
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection
Yanghao Li, Chao-Yuan Wu, Haoqi Fan +4
In this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection. We present an improved ve…