9 papers
OptiWorld: Optimal Control for Video World Generation under Physical Constraints
Yu Yuan, Jianhao Yuan, Xijun Wang +4
Video generation models are becoming a scalable form of world models, but they mainly generate plausible motion rather than proactively control or optimize the underlying dynamics.…
LVSA: Training-Free Sparse Attention for Long Video Diffusion
Gael Glorian, Ioannis Lamprou, Zhen Zhang +2
Dense self-attention is the compute and quality bottleneck of long-video diffusion inference: cost grows quadratically with the sequence length, and beyond the training horizon the…
Reflect to Inform: Boosting Multimodal Reasoning via Information-Gain-Driven Verification
Shuai Lv, Chang Liu, Feng Tang +5
Multimodal Large Language Models (MLLMs) achieve strong multimodal reasoning performance, yet we identify a recurring failure mode in long-form generation: as outputs grow longer,…
SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning
Xiuwei Chen, Wentao Hu, Hanhui Li +9
Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision…
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
Haoyuan Li, Yanpeng Zhou, Yufei Gao +7
Remarkable progress in 2D Vision-Language Models (VLMs) has spurred interest in extending them to 3D settings for tasks like 3D Question Answering, Dense Captioning, and Visual Gro…
SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning
Kun Xiang, Heng Li, Terry Jingchen Zhang +11
We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fu…