2 papers
cs.CV2024
Instance-free Text to Point Cloud Localization with Relative Position Awareness
Lichao Wang, Zhihao Yuan, Jinke Ren +2
Text-to-point-cloud cross-modal localization is an emerging vision-language task critical for future robot-human collaboration. It seeks to localize a position from a city-scale po…
cs.CV2024
Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding
Zhihao Yuan, Jinke Ren, Chun-Mei Feng +3
3D Visual Grounding (3DVG) aims at localizing 3D object based on textual descriptions. Conventional supervised methods for 3DVG often necessitate extensive annotations and a predef…