13 papers
LogiShot: Logically Coherent Cross-Shot Video Generation
Shuai Guo, Yuhang Yang, Zeyu Zhang +4
Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such as short-drama production,…
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
Chengjun Yu, Xuhan Zhu, Chaoqun Du +4
Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-ter…
Dual-Pathway Geometry-Aware MLLM for Spatial Intelligence
Yufei Zheng, Xuhan Zhu, Zide Liu +9
Spatial understanding of the physical world from 2D visual inputs hinges on two complementary forms of geometric knowledge: holistic 3D structural perception and fine-grained metri…
TIE: Time Interval Encoding for Video Generation over Events
Zhilei Shu, Shangwen Zhu, Zihang Liang +10
Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which 68% of general clips and ove…
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model
Chenfeng Wang, Wei He, Xuhan Zhu +10
In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likewise benefit from longer latent…
Stereo World Model: Camera-Guided Stereo Video Generation
Yang-Tian Sun, Zehuan Huang, Yifan Niu +4
We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or…