2 citations · 3 across the 4 of their papers we have counts for
1 paper · 2 filters
Yuanyuan Liu, Haiyang Mei, Dongyang Zhan +4
3D visual grounding (3DVG) identifies objects in 3D scenes from language descriptions. Existing zero-shot approaches leverage 2D vision-language models (VLMs) by converting 3D spat…