Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models
Xiaoyan Wang, Zeju Li, Yifan Xu +5
New era has unlocked exciting possibilities for extending Large Language Models (LLMs) to tackle 3D vision-language tasks. However, most existing 3D multimodal LLMs (MLLMs) rely on…
cs.CV2025
Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
Yifan Xu, Chao Zhang, Hanqi Jiang +6
Advancements in foundation models have made it possible to conduct applications in various downstream tasks. Especially, the new era has witnessed a remarkable capability to extend…
cs.CV2024
3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding
Zeju Li, Chao Zhang, Xiaoyan Wang +4
The remarkable potential of multi-modal large language models (MLLMs) in comprehending both vision and language information has been widely acknowledged. However, the scarcity of 3…