3d understanding 1multimodal large language models 1prior fusion 1spatial reasoning 1visual priors 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CV2026
Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding
Xiao Lin, Xiaohu Huang, Kai Han
The paper introduces ViPS, a framework that combines multiple visual priors from diverse foundation models using an Efficient Prior Proxy and Dynamic Prior Fusion to improve spatia…
cs.CV2026
PanoWorld: Real-World Panoramic Generation
Haoyuan Li, Dizhe Zhang, Yuemei Zhou +7
In this work, we aim to address the challenge of long-range memory in panoramic world models by exploiting the rotation-equivariant property of omnidirectional representations, whe…