works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CV2026

Surprise Forcing: What to Remember, When to Skip in Long Video Generation

Shuwei Shi, Zhen Li, Muyao Niu +4

Streaming autoregressive diffusion makes minute-scale video synthesis practical, but its bounded context and fixed denoising schedule allocate resources uniformly across a highly n…

cs.CV2026

From Pixels to States: Rethinking Interactive World Models as Game Engines

Zhen Li, Zian Meng, Shuwei Shi +4

The paper surveys how recent video generative models can be used to build interactive game worlds, analyzing four key dimensions of game engine design and introducing a large datas…

cs.CV2026

WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG

Zhen Li, Zian Meng, Shuwei Shi +5

Dynamical systems theory and reinforcement learning view world evolution as latent-state dynamics driven by actions, with visual observations providing partial information about th…

cs.CV2024

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation

Shuwei Shi, Biao Gong, Xi Chen +9

The image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware…

cs.CV2024

Mimir: Improving Video Diffusion Models for Precise Text Understanding

Shuai Tan, Biao Gong, Yutong Feng +6

Text serves as the key control signal in video generation due to its narrative nature. To render text descriptions into video clips, current video diffusion models borrow features…

cs.CV2024

ResMaster: Mastering High-Resolution Image Generation via Structural and Fine-Grained Guidance

Shuwei Shi, Wenbo Li, Yuechen Zhang +3

Diffusion models excel at producing high-quality images; however, scaling to higher resolutions, such as 4K, often results in over-smoothed content, structural distortions, and rep…