337 citations · 565 across the 11 of their papers we have counts for
5 papers · 1 filter
Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
Andy Zeng, Maria Attarian, Brian Ichter +10
Large pretrained (e.g., "foundation") models exhibit distinct capabilities depending on the domain of data they are trained on. While these domains are generic, they may only barel…
Grasping in the Wild:Learning 6DoF Closed-Loop Grasping from Low-Cost Demonstrations
Shuran Song, Andy Zeng, Johnny Lee +1
Intelligent manipulation benefits from the capacity to flexibly control an end-effector with high degrees of freedom (DoF) and dynamically react to the environment. However, due to…
ClearGrasp: 3D Shape Estimation of Transparent Objects for Manipulation
Shreeyak S. Sajjan, Matthew Moore, Mike Pan +4
Transparent objects are a common part of everyday life, yet they possess unique visual properties that make them incredibly difficult for standard 3D sensors to produce accurate de…
Im2Pano3D: Extrapolating 360 Structure and Semantics Beyond the Field of View
Shuran Song, Andy Zeng, Angel X. Chang +3
We present Im2Pano3D, a convolutional neural network that generates a dense prediction of 3D structure and a probability distribution of semantic labels for a full 360 panoramic vi…
Matterport3D: Learning from RGB-D Data in Indoor Environments
Angel Chang, Angela Dai, Thomas Funkhouser +6
Access to large, diverse RGB-D datasets is critical for training RGB-D scene understanding algorithms. However, existing datasets still cover only a limited number of views or a re…