9 papers
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
Weiliang Chen, Haowen Sun, Jun Gao +40
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, w…
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Hanyang Wang, Yimo Cai, Weiliang Chen +14
Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision-language models can recognize and explain diverse…
ParallelWorld: Test-Time Scaling for Embodied Reasoning
Min Chen, Shengjun Zhang, Yuxin Li +4
Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and interaction within physical environ…
Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models
Jindi Lv, Aoyu Li, Yuhao Zhou +6
Mamba demonstrates strong efficiency in modeling long visual sequences. However, when token reduction is applied to structurally enhanced Mamba variants, these models exhibit a sev…
TivTok: Broadcasting Time-Invariant Tokens for Scalable Video Tokenization
Weiliang Chen, Yuanhui Huang, Xuebo Wang +1
Video tokenization is fundamental to scalable video generation, as the number of tokens directly determines the computational cost and the length of videos that can be modeled. Exi…
MBench: A Comprehensive Benchmark on Memory Capability for Video World Models
Shengjun Zhang, Zhang Zhang, Simin Huang +11
Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelity visual sequences. However, a fundamental gap persists between…