12 citations · 13 across the 33 of their papers we have counts for
37 papers
Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
Yijun Yang, Shenghe Zheng, Wenbo Li +8
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue th…
Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation
Haoyu Wang, Songchun Zhang, Haoran Li +3
Action-conditioned video models require large-scale visual data paired with control signals that are temporally aligned with the resulting scene transitions. Such supervision is di…
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
Xionghao Wu, Yijun Yang, Shiyang Zhou +17
Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect…
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
Nan Duan, Haoyang Huang, Weiyang Jin +13
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain…
EchoWM: Open and Enterable Omnimodal World Models
Songchun Zhang, Yaowei Li, Junhao Zhuang +19
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music an…
HalluScope: Fine-grained Hallucination Diagnosis for Multimodal Large Language Models
Weilin Jin, Mingyu Wang, Wenbo Li +5
Although Multimodal Large Language Models have achieved strong performance across a wide range of vision-language tasks, they still suffer from hallucinations, where model outputs…