5 papers
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
Zhenghao Chen, Huiqun Wang, Di Huang
Multimodal large language models (MLLMs) are increasingly being applied to spatial cognition tasks, where they are expected to understand and interact with complex environments. Mo…
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
Yifan Li, Zhenghao Chen, Ziheng Wu +7
Recent advances in inference-time scaling, particularly those leveraging reinforcement learning with verifiable rewards, have substantially enhanced the reasoning capabilities of L…
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
Zhentao He, Can Zhang, Ziheng Wu +6
Recent advancements in multimodal large language models have enhanced document understanding by integrating textual and visual information. However, existing models exhibit incompl…
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
Yufei Zhan, Ziheng Wu, Yousong Zhu +10
Despite notable advancements in multimodal reasoning, leading Multimodal Large Language Models (MLLMs) still underperform on vision-centric multimodal reasoning tasks in general sc…
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Ziheng Wu, Zhenghao Chen, Ruipu Luo +6
Recently, vision-language models have made remarkable progress, demonstrating outstanding capabilities in various tasks such as image captioning and video understanding. We introdu…