3d scene understanding 1graph neural networks 1hierarchical attention 1multimodal large language models 1multi-room reasoning 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CV2026
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
He Liang, Chenyang Ma, Yiming Zhang +4
The paper introduces CAIRN, a topology‑aware large multimodal model that uses graph neural networks and hierarchical attention to understand and reason about multi‑room 3D scenes.
cs.CV2026
Scalable Visual Pretraining for Language Intelligence
Yiming Zhang, Zhonghan Zhao, Wenwei Zhang +14
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual…