activity
20242026
collaborators

13 papers

cs.CV2026

Improved Baselines with Representation Autoencoders

Jaskirat Singh, Boyang Zheng, Zongze Wu +3

Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders. In this paper, we systematically investigate several design choices and find three insigh…

cs.CV2025

What matters for Representation Alignment: Global Information or Spatial Structure?

Jaskirat Singh, Xingjian Leng, Zongze Wu +4

Representation alignment (REPA) guides generative training by distilling representations from a strong, pretrained vision encoder to intermediate diffusion features. We investigate…

cs.CV2025

From Slow Bidirectional to Fast Autoregressive Video Diffusion Models

Tianwei Yin, Qiang Zhang, Richard Zhang +4

Current video diffusion models achieve impressive generation quality but struggle in interactive applications due to bidirectional attention dependencies. The generation of a singl…

cs.CV2025

Long-Context State-Space Video World Models

Ryan Po, Yotam Nitzan, Richard Zhang +5

Video diffusion models have recently shown promise for world modeling through autoregressive frame prediction conditioned on actions. However, they struggle to maintain long-term m…

cs.CV2025

SliderSpace: Decomposing the Visual Capabilities of Diffusion Models

Rohit Gandikota, Zongze Wu, Richard Zhang +3

We present SliderSpace, a framework for automatically decomposing the visual capabilities of diffusion models into controllable and human-understandable directions. Unlike existing…

cs.CV2024

NewMove: Customizing text-to-video models with novel motions

Joanna Materzynska, Josef Sivic, Eli Shechtman +3

We introduce an approach for augmenting text-to-video generation models with customized motions, extending their capabilities beyond the motions depicted in the original training d…