Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
Siqu Ou, Tianrui Wan, Zhiyuan Zhao +2
While chain-of-thought (CoT) reasoning has substantially improved multimodal large language models (MLLMs) on complex reasoning tasks, existing approaches largely rely on long text…
cs.AI2025
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
Siqu Ou, Hongcheng Liu, Pingjie Wang +4
While chains-of-thought (CoT) have advanced complex reasoning in multimodal large language models (MLLMs), existing methods remain confined to text or static visual domains, often…