activity
20202026
most citedStyleSpace Analysis: Disentangled Controls for StyleGAN Image Generation

42 citations · 46 across the 10 of their papers we have counts for

collaborators

12 papers

cs.CV2026

EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing

Yuqian Zhou, Zhenghong Zhou, Zongze Wu +5

Interactive video generation and editing are becoming increasingly important for creative design. In this report, we introduce EditStream: a unified framework for interactive video…

cs.CV2026

Improved Baselines with Representation Autoencoders

Jaskirat Singh, Boyang Zheng, Zongze Wu +3

Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders. In this paper, we systematically investigate several design choices and find three insigh…

cs.CV2026

End-to-End Training for Unified Tokenization and Latent Denoising

Shivam Duggal, Xingjian Bai, Zongze Wu +5

Latent diffusion models (LDMs) enable high-fidelity synthesis by operating in learned latent spaces. However, training state-of-the-art LDMs requires complex staging: a tokenizer m…

cs.CV2026

Causality in Video Diffusers is Separable from Denoising

Xingjian Bai, Guande He, Zhengqi Li +3

Causality -- referring to temporal, uni-directional cause-effect relationships between components -- underlies many complex generative processes, including videos, language, and ro…

cs.CV2025

What matters for Representation Alignment: Global Information or Spatial Structure?

Jaskirat Singh, Xingjian Leng, Zongze Wu +4

Representation alignment (REPA) guides generative training by distilling representations from a strong, pretrained vision encoder to intermediate diffusion features. We investigate…

cs.CV2025

Improved Mean Flows: On the Challenges of Fastforward Generative Models

Zhengyang Geng, Yiyang Lu, Zongze Wu +3

MeanFlow (MF) has recently been established as a framework for one-step generative modeling. However, its ``fastforward'' nature introduces key challenges in both the training obje…