activity
20242026
collaborators

9 papers

cs.CV2026

OptiWorld: Optimal Control for Video World Generation under Physical Constraints

Yu Yuan, Jianhao Yuan, Xijun Wang +4

Video generation models are becoming a scalable form of world models, but they mainly generate plausible motion rather than proactively control or optimize the underlying dynamics.…

cs.CV2026

LVSA: Training-Free Sparse Attention for Long Video Diffusion

Gael Glorian, Ioannis Lamprou, Zhen Zhang +2

Dense self-attention is the compute and quality bottleneck of long-video diffusion inference: cost grows quadratically with the sequence length, and beyond the training horizon the…

cs.CV2026

Reflect to Inform: Boosting Multimodal Reasoning via Information-Gain-Driven Verification

Shuai Lv, Chang Liu, Feng Tang +5

Multimodal Large Language Models (MLLMs) achieve strong multimodal reasoning performance, yet we identify a recurring failure mode in long-form generation: as outputs grow longer,…

cs.CV2025

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning

Xiuwei Chen, Wentao Hu, Hanhui Li +9

Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision…

cs.CV2025

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs

Haoyuan Li, Yanpeng Zhou, Yufei Gao +7

Remarkable progress in 2D Vision-Language Models (VLMs) has spurred interest in extending them to 3D settings for tasks like 3D Question Answering, Dense Captioning, and Visual Gro…

cs.AI2025

SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning

Kun Xiang, Heng Li, Terry Jingchen Zhang +11

We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fu…