1 citations · 1 across the 3 of their papers we have counts for
7 papers · 1 filter
SAM 3D: 3Dfy Anything in Images
SAM 3D Team, Xingyu Chen, Fu-Jen Chu +20
We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single image. SAM 3D excels in natural images,…
Human-level 3D shape perception emerges from multi-view learning
Tyler Bonnen, Jitendra Malik, Angjoo Kanazawa
Humans can infer the three-dimensional structure of objects from two-dimensional visual inputs. Modeling this ability has been a longstanding goal for the science and engineering o…
Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
Vongani H. Maluleke, Kie Horiuchi, Lea Wilken +3
Understanding and generating multi-person interactions is a fundamental challenge with broad implications for robotics and social computing. While humans naturally coordinate in gr…
Reconstructing Hand-Held Objects in 3D from Images and Videos
Jane Wu, Georgios Pavlakos, Georgia Gkioxari +1
Objects manipulated by the hand (i.e., manipulanda) are particularly challenging to reconstruct from Internet videos. Not only does the hand occlude much of the object, but also th…
Reconstructing People, Places, and Cameras
Lea Müller, Hongsuk Choi, Anthony Zhang +3
We present "Humans and Structure from Motion" (HSfM), a method for jointly reconstructing multiple human meshes, scene point clouds, and camera parameters in a metric world coordin…
Estimating Body and Hand Motion in an Ego-sensed World
Brent Yi, Vickie Ye, Maya Zheng +6
We present EgoAllo, a system for human motion estimation from a head-mounted device. Using only egocentric SLAM poses and images, EgoAllo guides sampling from a conditional diffusi…