Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding
Xiao Lin, Xiaohu Huang, Kai Han
Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-t…
cs.CV2026
PanoWorld: Real-World Panoramic Generation
Haoyuan Li, Dizhe Zhang, Yuemei Zhou +7
In this work, we aim to address the challenge of long-range memory in panoramic world models by exploiting the rotation-equivariant property of omnidirectional representations, whe…