2 papers
cs.CL2025
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
Wenxin Zhu, Andong Chen, Yuchen Song +4
With the remarkable success of Multimodal Large Language Models (MLLMs) in perception tasks, enhancing their complex reasoning capabilities has emerged as a critical research focus…
cs.CV2025
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
Shaojun E, Yuchen Yang, Jiaheng Wu +3
In the latest advancements in multimodal learning, effectively addressing the spatial and semantic losses of visual data after encoding remains a critical challenge. This is becaus…