2 papers
cs.CV2025
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
Wencan Huang, Daizong Liu, Wei Hu
While 3D Multi-modal Large Language Models (MLLMs) demonstrate remarkable scene understanding capabilities, their practical deployment faces critical challenges due to computationa…
cs.CV2024
A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions
Daizong Liu, Yang Liu, Wencan Huang +1
Text-guided 3D visual grounding (T-3DVG), which aims to locate a specific object that semantically corresponds to a language query from a complicated 3D scene, has drawn increasing…