1 paper
Bo Gu, Zhikang Zhang, Zizhuang Wei +3
Recent multimodal large language models (MLLMs) have made remarkable progress in visual understanding and language-based reasoning, yet they lack a persistent world-centered repres…