2 papers
cs.CV2026
Scalable Visual Pretraining for Language Intelligence
Yiming Zhang, Zhonghan Zhao, Wenwei Zhang +14
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual…
cs.CV2026
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
He Liang, Chenyang Ma, Yiming Zhang +4
Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the ability to reason over real-world…