activity
20232026
most citedDyna-DepthFormer: Multi-frame Transformer for Self-Supervised Depth Estimation in Dynamic Scenes

2 citations · 4 across the 17 of their papers we have counts for

collaborators

20 papers

cs.CV2026

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

Zhuoran Zhao, Shengju Qian, Tongtong Liang +7

Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into cau…

cs.CV2026

Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

Haoyu Wang, Songchun Zhang, Haoran Li +3

Action-conditioned video models require large-scale visual data paired with control signals that are temporally aligned with the resulting scene transitions. Such supervision is di…

cs.CV2026

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Nan Duan, Haoyang Huang, Weiyang Jin +13

Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain…

cs.CV2026

EchoWM: Open and Enterable Omnimodal World Models

Songchun Zhang, Yaowei Li, Junhao Zhuang +19

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music an…

cs.CV2026

LiveLight: Real-time Streaming Video Relighting with Interactive Control

Yue Ma, Jiangming Wang, Yucheng Wang +8

We present LiveLight, the first diffusion-based framework for real-time streaming video relighting with interactive 3D lighting control. Achieving this is non-trivial, as it requir…

cs.CV2026

FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

Songchun Zhang, Sitong Guo, Xianghao Kong +4

Generative video compositing, which involves inserting external assets seamlessly into existing video sequences, is essential for content creation and visual effects. However, exis…