3 papers
cs.CV2026
MoWorld: A Flash World Model
Team Moxin, Deyi Ji, Tianrun Chen +37
The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive per…
cs.CV2026
HD-VGGT: High-Resolution Visual Geometry Transformer
Tianrun Chen, Yuanqi Hu, Yidong Han +11
High-resolution imagery is essential for accurate 3D reconstruction, as many geometric details only emerge at fine spatial scales. Recent feed-forward approaches, such as the Visua…
cs.RO2026
CReF: Cross-modal and Recurrent Fusion for Depth-conditioned Humanoid Locomotion
Yuan Hao, Ruiqi Yu, Shixin Luo +3
Stable traversal over geometrically complex terrain increasingly requires exteroceptive perception, yet prior perceptive humanoid locomotion methods often remain tied to explicit g…