337 citations · 780 across the 18 of their papers we have counts for
27 papers
UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding
Dave Zhenyu Chen, Ronghang Hu, Xinlei Chen +2
Performing 3D dense captioning and visual grounding requires a common and shared understanding of the underlying multimodal relationships. However, despite some previous attempts o…
Retrospectives on the Embodied AI Workshop
Matt Deitke, Dhruv Batra, Yonatan Bisk +36
We present a retrospective on the state of Embodied AI research. Our analysis focuses on 13 challenges presented at the Embodied AI Workshop at CVPR. These challenges are grouped i…
Understanding Pure CLIP Guidance for Voxel Grid NeRF Models
Han-Hung Lee, Angel X. Chang
We explore the task of text to 3D object generation using CLIP. Specifically, we use CLIP for guidance without access to any datasets, a setting we refer to as pure CLIP guidance.…
Articulated 3D Human-Object Interactions from RGB Videos: An Empirical Analysis of Approaches and Challenges
Sanjay Haresh, Xiaohao Sun, Hanxiao Jiang +2
Human-object interactions with articulated objects are common in everyday life. Despite much progress in single-view 3D reconstruction, it is still challenging to infer an articula…
OPD: Single-view 3D Openable Part Detection
Hanxiao Jiang, Yongsen Mao, Manolis Savva +1
We address the task of predicting what parts of an object can open and how they move when they do so. The input is a single image of an object, and as output we detect what parts o…
Interpretation of Emergent Communication in Heterogeneous Collaborative Embodied Agents
Shivansh Patel, Saim Wani, Unnat Jain +4
Communication between embodied AI agents has received increasing attention in recent years. Despite its use, it is still unclear whether the learned communication is interpretable…