1 paper · 1 filter
Austin T. Wang, ZeMing Gong, Angel X. Chang
3D visual grounding (3DVG) involves localizing entities in a 3D scene referred to by natural language text. Such models are useful for embodied AI and scene retrieval applications,…