3d scene understanding 1graph neural networks 1hierarchical attention 1multimodal large language models 1multi-room reasoning 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
He Liang, Chenyang Ma, Yiming Zhang +4
The paper introduces CAIRN, a topology‑aware large multimodal model that uses graph neural networks and hierarchical attention to understand and reason about multi‑room 3D scenes.
cs.LG2026
RiTTA: Modeling Event Relations in Text-to-Audio Generation
Yuhang He, Yash Jain, Xubo Liu +2
Despite significant advancements in Text-to-Audio (TTA) generation models achieving high-fidelity audio with fine-grained context understanding, they struggle to model the relation…
cs.SD2024
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
Yuhang He, Sangyun Shin, Anoop Cherian +2
Accurately localizing 3D sound sources and estimating their semantic labels -- where the sources may not be visible, but are assumed to lie on the physical surface of objects in th…