2 papers
cs.CV2025
Spatial Preference Rewarding for MLLMs Spatial Understanding
Han Qiu, Peng Gao, Lewei Lu +3
Multimodal large language models~(MLLMs) have demonstrated promising spatial understanding capabilities, such as referencing and grounding object descriptions. Despite their succes…
cs.CV2025
Multimodal 3D Reasoning Segmentation with Complex Scenes
Xueying Jiang, Lewei Lu, Ling Shao +1
The recent development in multimodal learning has greatly advanced the research in 3D scene understanding in various real-world tasks such as embodied AI. However, most existing st…