4 papers · 1 filter
SeqAlign3DVG: A Sequence-Aligned Benchmark and Voxel Reasoning Framework for 3D Visual Grounding
Yi Zhang, Yi Wang, Yueting Wu +3
Image-based 3D visual grounding is critical for embodied agents, yet existing benchmarks suffer from loose text-observation alignment and neglect temporal ordering. We introduce Se…
GVSynergy-Det: Synergistic Gaussian-Voxel Representations for Multi-View 3D Object Detection
Yi Zhang, Yi Wang, Lei Yao +1
Image-based 3D object detection aims to identify and localize objects in 3D space using only RGB images, eliminating the need for expensive depth sensors required by point cloud-ba…
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
Lei Yao, Yi Wang, Yi Zhang +2
The significance of informative and robust point representations has been widely acknowledged for 3D scene understanding. Despite existing self-supervised pre-training counterparts…
3DGeoDet: General-purpose Geometry-aware Image-based 3D Object Detection
Yi Zhang, Yi Wang, Yawen Cui +1
This paper proposes 3DGeoDet, a novel geometry-aware 3D object detection approach that effectively handles single- and multi-view RGB images in indoor and outdoor environments, sho…