1 paper
Yuzhen Li, Min Liu, Zhaoyang Li +4
Monocular 3D Visual Grounding (Mono3DVG) is an emerging task that locates 3D objects in RGB images using text descriptions with geometric cues. However, existing methods face two k…