3d scene understanding 1graph neural networks 1hierarchical attention 1multimodal large language models 1multi-room reasoning 1
From the 1 of 11 linked papers with an AI index.
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
He Liang, Chenyang Ma, Yiming Zhang +4
The paper introduces CAIRN, a topology‑aware large multimodal model that uses graph neural networks and hierarchical attention to understand and reason about multi‑room 3D scenes.
cs.CV2024
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
Chenyang Ma, Kai Lu, Ta-Ying Cheng +2
Current state-of-the-art spatial reasoning-enhanced VLMs are trained to excel at spatial visual question answering (VQA). However, we believe that higher-level 3D-aware tasks, such…