1 paper
Henry Zheng, Chenyue Fang, Rui Huang +3
Vision-language models (VLMs) have achieved strong performance in multimodal understanding and reasoning, yet grounded reasoning in 3D scenes remains underexplored. Effective 3D re…