1 paper
Alaa Asfour, Christopher Indris, Leihan Chen +2
Large-scale 3D vision-language models (VLMs) like LLaVA-3D offer strong spatial reasoning but are difficult to deploy due to high computational costs. We propose a knowledge distil…