1 paper
Chenshu Hou, Liang Peng, Xiaopei Wu +2
3D visual grounding aims to identify objects in 3D point cloud scenes that match specific natural language descriptions. This requires the model to not only focus on the target obj…