18 papers
Masked Visual Actions for Unified World Modeling
Hadi Alzayer, Wenlong Huang, Haonan Chen +8
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challe…
TinyHistory: Lightweight Video History Embeddings via Two-Stage Context Learning
Lvmin Zhang, Shengqu Cai, Muyang Li +6
History context is central to autoregressive video generation, driving consistency and storytelling for both commercial models and personal use cases. For example, personal users,…
Spectral Progressive Diffusion for Efficient Image and Video Generation
Howard Xiao, Brian Chao, Lior Yariv +1
Diffusion models have been shown to implicitly generate visual content autoregressively in the frequency domain, where low-frequency components are generated earlier in the denoisi…
GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation
Jan Ackermann, Shengqu Cai, Boyang Deng +3
Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only implicitly, leading to object def…
WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform
Yu Shang, Yinzhou Tang, Yiding Ma +22
World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, ex…
MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines
Ryan Po, David Junhao Zhang, Amir Hertz +3
Video world models have shown immense promise for interactive simulation and entertainment, but current systems still struggle with two important aspects of interactivity: user con…