collaborators

21 papers

cs.CV2026

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Nan Duan, Haoyang Huang, Weiyang Jin +13

Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain…

cs.CV2026

EchoWM: Open and Enterable Omnimodal World Models

Songchun Zhang, Yaowei Li, Junhao Zhuang +19

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music an…

cs.CV2026

Self Gradient Forcing: Native Long Video Extrapolation

Junhao Zhuang, Shiyi Zhang, Yuxuan Bian +11

Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-tru…

cs.CV2026

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization

Xiaoxuan He, Siming Fu, Zeyue Xue +9

Group Relative Policy Optimization has emerged as essential for aligning video diffusion models with human preferences, but faces a critical computational bottleneck: training a 14…

cs.CV2026

Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions

Luxury, Jie Huang, Zihao Fan +25

While recent autoregressive video diffusion models achieve remarkable streaming quality, they remain confined to low resolutions (e.g., 480P), leaving efficient, scalable, real-tim…

cs.CV2026

HPSv3++: Scaling Reward Models Across the Full Spectrum of Diffusion Model Capabilities

Yijun Liu, Jie Huang, Zeyue Xue +5

Reward models guide text-to-image (T2I) systems toward outputs aligned with human preferences. However, typical reward models such as HPSv3 are trained on pre-annotated data from e…