Zhenxiang Lin, Xidong Peng, Peishan Cong +6
We introduce the task of 3D visual grounding in large-scale dynamic scenes based on natural linguistic descriptions and online captured multi-modal visual data, including 2D images…