collaborators

6 papers

cs.CV2026

Astra: General Interactive World Model with Autoregressive Denoising

Yixuan Zhu, Jiaqi Feng, Wenzhao Zheng +5

Recent advances in diffusion transformers have empowered video generation models to generate high-quality video clips from texts or images. However, world models with the ability t…

cs.CV2025

MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives

Sihui Ji, Xi Chen, Shuai Yang +3

The core challenge for streaming video generation is maintaining the content consistency in long context, which poses high requirement for the memory design. Most existing solution…

cs.CV2025

UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation

Jiehui Huang, Yuechen Zhang, Xu He +7

Recent video generation models demonstrate impressive synthesis capabilities but remain limited by single-modality conditioning, constraining their holistic world understanding. Th…

cs.CV2025

Terra: Explorable Native 3D World Model with Point Latents

Yuanhui Huang, Weiliang Chen, Wenzhao Zheng +4

World models have garnered increasing attention for comprehensive modeling of the real world. However, most existing methods still rely on pixel-aligned representations as the basi…

cs.CV2025

PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning

Sihui Ji, Xi Chen, Xin Tao +2

Video generation models nowadays are capable of generating visually realistic videos, but often fail to adhere to physical laws, limiting their ability to generate physically plaus…

cs.LG2025

Mitigating the Noise Shift for Denoising Generative Models via Noise Awareness Guidance

Jincheng Zhong, Boyuan Jiang, Xin Tao +3

Existing denoising generative models rely on solving discretized reverse-time SDEs or ODEs. In this paper, we identify a long-overlooked yet pervasive issue in this family of model…