2 papers
cs.CV2025
AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models
Xinyi Wang, Xun Yang, Yanlong Xu +3
Effective human-agent collaboration in physical environments requires understanding not only what to act upon, but also where the actionable elements are and how to interact with t…
cs.CV2025
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
Zhihao Yuan, Shuyi Jiang, Chun-Mei Feng +4
Currently, utilizing large language models to understand the 3D world is becoming popular. Yet existing 3D-aware LLMs act as black boxes: they output bounding boxes or textual answ…