3 papers
cs.CV2025
Details Matter for Indoor Open-vocabulary 3D Instance Segmentation
Sanghun Jung, Jingjing Zheng, Ke Zhang +10
Unlike closed-vocabulary 3D instance segmentation that is often trained end-to-end, open-vocabulary 3D instance segmentation (OV-3DIS) often leverages vision-language models (VLMs)…
cs.CV2025
PointLAMA: Latent Attention meets Mamba for Efficient Point Cloud Pretraining
Xuanyu Lin, Xiaona Zeng, Xianwei Zheng +1
Mamba has recently gained widespread attention as a backbone model for point cloud modeling, leveraging a state-space architecture that enables efficient global sequence modeling w…
cs.GR2025
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation
Liu He, Xiao Zeng, Yizhi Song +7
Multimodal Large Language Models (MLLMs) struggle with accurately capturing camera-object relations, especially for object orientation, camera viewpoint, and camera shots. This ste…