9 papers
SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning
Xiuwei Chen, Wentao Hu, Hanhui Li +9
Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision…
OptiWorld: Optimal Control for Video World Generation under Physical Constraints
Yu Yuan, Jianhao Yuan, Xijun Wang +4
Video generation models are becoming a scalable form of world models, but they mainly generate plausible motion rather than proactively control or optimize the underlying dynamics.…
LVSA: Training-Free Sparse Attention for Long Video Diffusion
Gael Glorian, Ioannis Lamprou, Zhen Zhang +2
Dense self-attention is the compute and quality bottleneck of long-video diffusion inference: cost grows quadratically with the sequence length, and beyond the training horizon the…
Reflect to Inform: Boosting Multimodal Reasoning via Information-Gain-Driven Verification
Shuai Lv, Chang Liu, Feng Tang +5
Multimodal Large Language Models (MLLMs) achieve strong multimodal reasoning performance, yet we identify a recurring failure mode in long-form generation: as outputs grow longer,…
From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
Jiahui Zhang, Yurui Chen, Yanpeng Zhou +10
Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unl…
AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning
Kun Xiang, Zhili Liu, Terry Jingchen Zhang +12
In this paper, we address the challenging task of multimodal reasoning by incorporating the notion of ``slow thinking'' into multimodal large language models (MLLMs). Our core idea…