4 papers
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
Jiaxin Huang, Ziwen Li, Hanlve Zhang +6
The integration of language and 3D perception is critical for embodied AI and robotic systems to perceive, understand, and interact with the physical world. Spatial reasoning, a ke…
MMGen: Unified Multi-modal Image Generation and Understanding in One Go
Jiepeng Wang, Zhaoqing Wang, Hao Pan +4
A unified diffusion framework for multi-modal generation and understanding has the transformative potential to achieve seamless and controllable image diffusion and other cross-mod…
PanoSLAM: Panoptic 3D Scene Reconstruction via Gaussian SLAM
Runnan Chen, Zhaoqing Wang, Jiepeng Wang +4
Understanding geometric, semantic, and instance information in 3D scenes from sequential video data is essential for applications in robotics and augmented reality. However, existi…
OVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies
Runnan Chen, Xiangyu Sun, Zhaoqing Wang +8
Open-vocabulary scene understanding using 3D Gaussian (3DGS) representations has garnered considerable attention. However, existing methods mostly lift knowledge from large 2D visi…